SmartGateSmartGate

OWASP LLM Top 10: How to Read It and What to Fix First

The OWASP LLM Top 10 is a list of ten risk categories, ordered by how common and how damaging each one is when large language models are used in an application. It is a reading aid rather than a control catalogue: some entries are design problems you can only build your way out of, some are configuration you can switch on this week, and some are supplier risk you mostly manage with contracts and…

Short answer: The OWASP LLM Top 10 is a list of ten risk categories, ordered by how common and how damaging each one is when large language models are used in an application. It is a reading aid rather than a control catalogue: some entries are design problems you can only build your way out of, some are configuration you can switch on this week, and some are supplier risk you mostly manage with contracts and version pinning. Read it as an index, map each entry to the kind of control it needs, and land it in a fixed order — starting with what the model is allowed to do, not with the most famous entry.

Key takeaways

  • The list names risks, not fixes. Each entry is a category with its own page; none of them is a single control you install, and the numbering is a rough priority, not a work order.
  • Sort the ten into three buckets before you do anything else — design, configuration, and supply and dependency risk — because the bucket decides who owns the fix and how fast it can land.
  • The highest-value first move is an authority inventory: what the model can read, what it can call, and where it can send data. That one map shrinks the blast radius of most of the list.
  • Configuration beats procurement for the first three controls. Least privilege, a confirmation step in front of mutating actions, and treating model output as untrusted cost no new product.
  • Use it as a checklist with evidence, not a posture. One row per entry, one question, one owner, one artefact — otherwise it becomes the annual document nobody opens.

What the OWASP LLM Top 10 is, and what it is not

The OWASP Top 10 for LLM Applications is a community project, not a standard with an audit behind it. OWASP describes it as a guide to the most critical security risks facing applications powered by large language models, built from the project's own research and the incidents its contributors see in the field. The public index lists ten entries, each with its own page that states the risk, gives example scenarios and collects mitigations. The numbering runs LLM01 to LLM10 and is the project's judgement of relative importance, not a sequence to work through (OWASP Top 10 for LLM Applications).

It is worth being precise about what the list is for, because two mistakes are common. The first is treating it as a compliance target: there is no certificate, no auditor and no pass or fail. The second is treating it as a testing plan: an entry such as Excessive Agency is a property of how you built the system, and a scanner cannot decide whether your system has it. What the list actually gives you is a shared vocabulary — ten names that an engineer, a security reviewer and a product owner can all use to point at the same thing in a design review.

The list also covers more ground than the phrase "prompt injection" suggests, which is the point of reading it as a whole. Some entries are about the model's input, some about its output, some about what it is allowed to do, and some about the components around it: the data it was trained on, the vectors it retrieves from, the packages it was built with. A team that fixes only the most famous entry has fixed one input path and left the rest.

The ten entries, indexed by the kind of control they answer

Here is the whole list in one place, with the official title of each entry and a plain reading of the kind of control it needs. The middle column is the useful one: it tells you who owns the fix and whether the fix is a build change, a setting, or a supplier conversation. The last column points at the cluster page that goes deeper for that risk, so this page can stay an index.

# Entry (official title) Kind of control it answers Where this cluster goes deeper
LLM01 Prompt Injection Design — the trust boundary between instructions and data prompt injection in depth
LLM02 Sensitive Information Disclosure Design plus data governance — what the model can reach Framework mapping below
LLM03 Supply Chain Supply — provenance and version control of dependencies the model supply chain
LLM04 Data and Model Poisoning Supply — integrity of the data and weights you did not make Model and dataset provenance
LLM05 Improper Output Handling Design — output is untrusted input to the next system the guardrail layer
LLM06 Excessive Agency Configuration — the permissions and tools you granted the sandbox around tool execution
LLM07 System Prompt Leakage Configuration — what you put in the prompt and how it is handled Prompt and secret hygiene
LLM08 Vector and Embedding Weaknesses Supply plus configuration — the retrieval store and its access Retrieval and access control
LLM09 Misinformation Design plus evaluation — what the output is allowed to assert Evaluation and grounding
LLM10 Unbounded Consumption Configuration — the limits and budgets on calls and loops Metering and rate limits

Read down the middle column and the shape of the list appears. Four entries are design problems, so they are fixed by changing how the application is assembled — the boundary between trusted instructions and untrusted data, what the model is allowed to reach, whether its output is trusted by the next system, and what it is allowed to assert. Five are configuration or data governance problems, so they are fixed by changing settings, permissions and stored content. Three touch supply: the model and its packages, the data and weights behind it, and the retrieval store beside it. Two entries sit in more than one bucket, and that overlap is real rather than sloppy — supply chain problems become design problems the moment you ship the dependency.

Reading the list in three buckets: design, configuration, supply

The three buckets are not a re-ranking; they are the fastest way to route the work. A design problem cannot be closed by a setting, and a configuration problem should not wait for a redesign. Sorting first prevents the common failure of starting an architectural rewrite to answer a question that a policy change would have resolved.

Design problems are the entries where the harm is created by how the pieces fit together, so the remedy is a change to the architecture or the data flow: prompt injection, sensitive information disclosure, improper output handling, and the grounding half of misinformation. You cannot configure your way to a trust boundary. The work belongs to the engineers who assemble the context window and the tool calls, and it moves at the speed of a code change.

Configuration problems are the entries where the mechanism already exists and the question is whether you switched it on: excessive agency, system prompt leakage, unbounded consumption, and much of vector and embedding weakness. Least privilege, a human confirmation step, a hard budget, and a rule about what may be written into memory all live here. The work belongs to whoever owns the platform, and a meaningful share of it can land in a week.

Supply and dependency problems are the entries where you are trusting something you did not build: the model, the training data, the weights, the packages, and the retrieval store. You manage them with provenance, version pinning, integrity checks and supplier questions rather than with code you write yourself. The work belongs to whoever signs the vendor agreements, and it moves at the speed of procurement — which is exactly why it gets deferred, and why it should be scheduled.

The buckets also predict where a control will be weak. A design control is only as good as the engineer who implements it, so it needs review. A configuration control is only as good as its default, so it needs a test that fails when the setting is off. A supply control is only as good as your visibility into the supplier, so it needs a documented assumption that someone re-checks. The sibling framework-to-control mapping connects the same list to the governance frameworks, which is the right next step once the buckets are understood.

OWASP agentic AI top 10: the companion list for agents that act

The LLM Top 10 was written for applications where a model reads and writes text. The moment the model plans and calls tools — the agentic case — OWASP maintains a second, separate list for it: the OWASP Top 10 for Agentic Applications, published for 2026, which addresses autonomous systems that act across workflows rather than single model calls. It is a distinct index with its own entry numbering, and it should be read beside the LLM list rather than instead of it (OWASP Top 10 for Agentic Applications).

The practical reading is that the LLM list is the substrate and the agentic list is the delta. An agent inherits every LLM risk and adds risks that only exist because it can act: goal manipulation across a long horizon, degraded memory, unsafe tool composition, and failures of identity and delegation when one agent acts for another. That is why the agentic list is organised around the agent's lifecycle — how it is discovered, how it is authenticated, how it is sandboxed, how its tools are scoped — rather than around the input and output of a single call.

For a small team the implication is a sequencing one. If you are building a chat-style assistant, the LLM list is the one to read. If you are building anything that calls tools on its own, both lists apply, and the agentic entries raise the cost of the configuration bucket: least privilege and a hard budget stop being hygiene and become the controls that bound what an autonomous loop can do before anyone notices it. The same three buckets sort either list, because design, configuration and supply are properties of the system rather than of the model.

OWASP LLM Top 10 2025, the 2023-24 list, and the moving editions

Readers arrive with different years attached to the phrase, so it helps to know how the project keeps time. The entries published under the official index carry 2025 in their identifiers — LLM01:2025 through LLM10:2025 — and that set is what people mean today when they say "the OWASP LLM Top 10". The project also keeps an archive of the earlier 2023-24 list, which had a different shape and fewer entries, so a reader comparing two blog posts may be comparing two editions without knowing it (the 2023-24 list).

The editions move for a good reason rather than a cosmetic one. The 2023-24 list was assembled from early research on a technology that was, at the time, mostly a chat product. The 2025 set was re-weighted by what the community had seen in production, which is why supply chain and data poisoning appear as prominent entries and why "excessive agency" exists as its own category at all. The project has since published a newer edition of the LLM list, so the specific ranking and wording can move again between the edition you read and the one your supplier quotes (the OWASP GenAI project index).

The practical consequence is a reading habit rather than a fact to memorise: cite the edition and the entry number, not just the headline. "We handle LLM01" means something specific inside the 2025 edition and can shift if the numbering is re-cut in a later one. When a stakeholder asks whether you cover the list, the honest answer names the edition, names the entries, and states which are in scope for your system — because a risk index that is applied to everything is applied to nothing in practice.

OWASP LLM Top 10 GitHub, and where the canonical text lives

The canonical home of the list is the OWASP GenAI Security Project site, and each entry has its own page under the project's llmrisk path with the risk statement, scenarios and mitigations. That is the version to cite, because it is the one the project maintains. The project's source material and the broader list initiative also live in the open on GitHub, which is where the work is drafted and reviewed and where the community discusses changes before they land (the OWASP LLM Top 10 project on GitHub).

For a team that is actually going to use the list, the GitHub home matters for a narrower reason than provenance: it is where you can see a change coming. Because the list is community-edited, an entry can be reworded or re-weighted between editions, and a team that has mapped its controls to entry numbers wants to know when those numbers move. Watching the project is a cheap way to avoid finding out from a supplier's marketing page that the edition you built your review around has been superseded.

The list is a starting point, not a complete model of your risk. It is a general index for LLM applications, and your system has risks that are specific to its domain — a marketplace has fraud risks the list will not name, a healthcare tool has privacy obligations it only gestures at. Use the OWASP list to cover the common ground, and your own threat model for the rest. The list's value is that it stops the common ground from being forgotten, not that it replaces the local analysis.

The priority order for a small team: the first three controls

A small team cannot land ten entries at once, and it should not try. Three controls, done in order, shrink the blast radius of almost the whole list and cost nothing to buy. They are all in the design and configuration buckets, which is deliberate: they are the entries where a change you make this quarter changes the consequence of a failure, not just the wording of a document.

First, build the authority inventory. Write down three lists: what the model can read (the data sources and retrievals it is given), what it can call (every tool and every action, including the ones that only read), and where it can send data outward. This inventory is the practical answer to Excessive Agency, and it also tells you how exposed Sensitive Information Disclosure and Prompt Injection actually are, because those risks only bite where the three lists meet. A system where the three lists do not intersect is a system where the worst entries cannot cause damage. Do this before you buy anything.

Second, cut the outbound leg and add a confirmation step. From the inventory, remove the tools and destinations the task does not need, and put a human confirmation in front of every action that writes, pays, deletes or sends. This is the single highest-value configuration change on the list, because it does not depend on recognising a malicious input — it bounds what any input can cause. It answers the Excessive Agency entry directly and it is the control that makes Prompt Injection survivable rather than preventable, which is the honest framing: an injected instruction with nowhere to send what it gathered is a bad answer, not a breach.

Third, treat the model's output as untrusted, and cap the loop. Output is the input to whatever consumes it — a shell, a database, a browser, a template — and Improper Output Handling is the entry that exists because teams forget this. Validate and escape model output at the boundary where it meets another system, and never let it decide a privileged action on its own. At the same time, put a hard limit on tokens, steps and wall-clock time per task, which is the cheapest way to answer Unbounded Consumption and to bound a runaway loop before it becomes an invoice.

Everything else on the list is real and can be scheduled after these three. Supply chain and data poisoning are supplier conversations with their own timeline. System prompt leakage, vector and embedding weaknesses, and misinformation each have a specific control, and none of them is the first thing to build. The point of the ordering is not that the others do not matter; it is that the first three change what the others can cost you, so they should not wait behind a procurement cycle.

Turning the top 10 into a checklist you can actually run

The list becomes useful the moment it stops being a document and starts being a table with an owner and an artefact per row. The format that survives contact with a real team is boring on purpose: one row per entry, a question you can answer yes or no, the artefact that answers it, and a named owner. If a row cannot be answered with an artefact — a config file, a log, a contract — the row is a wish, not a control.

Entry The question to ask The artefact that answers it
LLM01 Prompt Injection Can untrusted text reach a tool that acts? The authority inventory, plus the confirmation policy
LLM02 Sensitive Information Disclosure What data can the model read, and is all of it needed? A data-access list with a justification per source
LLM03 Supply Chain Do we know every model and package version we run? A pinned dependency and model manifest
LLM04 Data and Model Poisoning Where did the training and fine-tuning data come from? A provenance record with the supplier named
LLM05 Improper Output Handling Is model output escaped before the next system reads it? The validation code at the output boundary
LLM06 Excessive Agency Does every tool have a permission it actually needs? The tool list, with a least-privilege review date
LLM07 System Prompt Leakage Does the prompt hold anything that is a secret? A prompt template with no credentials in it
LLM08 Vector and Embedding Weaknesses Who can write to the retrieval store, and who can read it? Access rules on the index, and a write-review step
LLM09 Misinformation What must the output never assert without a source? The evaluation cases that check grounding
LLM10 Unbounded Consumption Is there a hard cap on tokens, steps and time? The budget and rate settings, with a test that fails when off

Two habits keep the table alive rather than archived. The first is to attach every row to a system you can point at, so the checklist names your agent rather than a concept; a row that applies to "LLMs in general" is a row nobody owns. The second is to review the table when the edition moves, because the entry numbering is the one thing about this list that is genuinely unstable, and a mapping to a superseded edition is worse than no mapping. The adversarial half of the review — the part where someone tries to make the controls fail — is a separate discipline, and red-teaming as a method is where this cluster covers it.

Where SmartGate fits

SmartGate does not detect or block any entry on the OWASP list, and this page will not imply otherwise. What it provides maps to the configuration bucket, and to two entries in particular: the metering and limits that answer Unbounded Consumption, and the recorded call history that makes the authority inventory answerable with evidence rather than memory. SmartGate is an MCP-native algorithm gateway for token control, traffic shaping and agent audit — a single authenticated endpoint through which an agent reaches its tools, with per-key metering and an audit record written as calls happen.

That matters to this list for a specific and modest reason. Two of the ten entries — Excessive Agency and Unbounded Consumption — are about what the agent is permitted to do and how much it may do, and both are easier to answer when permissions and limits live at one enforcement point outside the prompt rather than in the system message. A tool call that passes through the gateway can be counted against a key, capped by a budget, and written to a record that a reviewer can read later; the same call described in a prompt is not a control. The surfaces are documented rather than described here: call control on hard budget settings, the record on the audit trail, and the plan limits that decide retention and rate. Seven tools are exposed through the gateway — smart_fetch, smart_search, smart_context_gate, smart_dedup, smart_budget_guard, smart_memory and smart_pipe — and each call is counted against the caller's key. The pricing page is the authoritative table for monthly token caps, per-key request rates and audit-log retention; treat the tiers as operational limits rather than as a claim about which risks they solve.

OWASP LLM Top 10 vulnerabilities, mapped to our control points

Every section above maps the official list onto a kind of control. This one is the only first-hand one: it reads the control plane that serves this site and names where each mechanism lives in the source, so a reader can check it rather than take it on trust. The reading is of origin/main, on 2026-10-08, and it is a read of what the code enforces — not a claim that the gateway detects any OWASP category, which it does not.

Unbounded Consumption and Excessive Agency, as limits rather than policy. The two entries about what an agent may do and how much it may do are the ones a gateway can actually answer, because both are counts. backend/smartgate/core/rate_limiter.py exposes check_rate_limit_mcp(key_id, team_id, *, per_key_limit, team_ceiling, window_seconds=60): one request increments two fixed-window counters in the same 60-second bucket — rate_limit_mcp:<key_id>:<bucket> for the single key and rate_limit_mcp_team:<team_id>:<bucket> for the whole team — and the refusal names which dimension stopped it, limit_scope being mcp_key or mcp_team, with a retry_after hint. A REST-side team limiter, check_rate_limit, carries the same shape with limit_scope rest_team. Because the two counters are separate, one busy client cannot spend a team's whole minute, and one team cannot spend its way past a single key's ceiling.

The numbers are a table, not a preference. backend/smartgate/core/plan_entitlements.py pins a CATALOG_RATE_LIMITS dict with three fields per tier — rest_write_rpm, mcp_rpm_per_key, mcp_rpm_team_ceiling — and merge_rate_limits(plan, features) lets a team's stored feature row override the catalog default, cached by resolve_rate_limits(team_id) for 60 seconds. The limit a call meets is therefore a configured value with a named origin, not something a caller can argue with at run time; the operational per-key and monthly figures a customer sees remain the ones on the pricing page.

The credential surface is a digest. backend/smartgate/core/api_key_hash.py stores a key as hash_api_key(raw_key), a SHA-256 digest over the salt, a separator and the raw key, hex-encoded, and the module's own contract is that it matches the Next app's lib/api-keys hashApiKey — so the value that identifies a caller is one-way rather than the key itself, and the two sides of the stack compute it the same way. It is secret hygiene at the credential, and it is not a control for the content of LLM02 or LLM07.

The record is exportable, which is what makes the authority inventory answerable. The inventory this page tells you to write is only as good as the call log behind it. backend/smartgate/core/audit_export.py projects each audit row to a fixed SIEM_FIELDS tuple — timestamp, request_id, correlation_id, trace_id, agent_platform, route, transport, key_id, source_ip, tool, token_used, latency_ms, success, params — and stream_export emits them as NDJSON or CSV, with scope one of core, access or full, a window (since) and a limit clamped to MAX_EXPORT_LIMIT of 50,000 rows. The export is quota'd too: check_export_rate_limit allows 10 exports per hour per team. Those field names are the artefact for LLM06 (what was actually called, by whose key) and the reason the authority inventory can be answered with evidence rather than memory.

What this does not claim. None of the four mechanisms reads the content of a request, and none of them maps to a prompt-injection or data-poisoning control. They are configuration-bucket answers — metering, a two-dimensional rate cap, a one-way key digest and an exportable audit row — and they sit beside, not in place of, the design controls this page tells you to build first.

Frequently Asked Questions

Is the OWASP LLM Top 10 a compliance standard?

No. It is a community risk index with no certification and no auditor, so there is nothing to pass. It is useful as a shared vocabulary and a checklist, and the accountability for the decisions it lists stays with your team rather than transferring to OWASP.

Which entry should we fix first?

Start with the authority inventory, not with the most famous entry. Writing down what the model can read, what it can call and where it can send data answers Excessive Agency directly and tells you how exposed the rest of the list is. Then cut the outbound leg and add a confirmation step before any mutating action.

How is the agentic list different from the LLM list?

The LLM list was written for applications where a model reads and writes text. The agentic list was written for systems that plan and call tools, so it adds risks that only exist once the agent can act, such as goal manipulation across many steps and unsafe delegation between agents. If your system acts on its own, both lists apply.

Does the numbering change between editions?

Yes, and that is the main reason to cite the edition and the entry number together rather than the headline alone. The 2023-24 list had a different shape from the 2025 set, and the project has since published a newer edition. A control mapped to a superseded numbering means less than a control mapped to the current one.

Can one product close the whole list?

No, and a vendor that says it can should be read carefully. The ten entries span design, configuration, and supplier risk; a single product can sit in one or two of those buckets at most. Design problems are closed by how you assemble the application, and supply problems are closed by what you know about what you did not build.

Limitations

This page is a reading of the list, not a substitute for it, and it deliberately stops at the index. It does not restate the attack scenarios or the per-entry mitigations, because those belong to the OWASP entry pages and to this cluster's deeper sibling pages, and quoting them here would make a shorter version of documents the reader can open directly. Where the list is silent, this page is silent too.

Two honest caveats follow. The bucketing into design, configuration and supply is a reading aid this page proposes; OWASP does not sort its own entries that way, and a specific system can put an entry in a different bucket than the one used here. And the priority order is a starting point for a small team, not a claim about every organisation: a regulated deployment may have to schedule supply chain and data provenance first because a contract or an audit requires it, whatever the blast-radius argument says. The edition and entry numbers move, so the page states its sources and their retrieval date rather than freezing a ranking that the project itself revises.

Sources

Method note

This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned 0 of 7 sections for this page (0 abstention(s), 7 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the seven section keywords, and the remote candidate fallback returned generic helpers that are collisions rather than section-specific evidence about a published OWASP list. A pinned generic would have given the page the shape of a verified article with none of the substance, so every section above is written from the official OWASP sources named beside it. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.