What Is Prompt Injection? Attack Surface and Defenses
Prompt injection is the class of attack in which text the developer did not write — a user message, a fetched web page, an email, a document, or a tool's return value — is read by a language model as an instruction, so the model acts on someone else's intent inside your application.
Short answer: Prompt injection is the class of attack in which text the developer did not write — a user message, a fetched web page, an email, a document, or a tool's return value — is read by a language model as an instruction, so the model acts on someone else's intent inside your application. It stops being a novelty the moment the model can act: once the same context that reads attacker-influenced content can also call tools, reach private data and send data out, injected text is no longer bad output but an instruction path with your application's authority behind it. This page is the cluster's centre: it defines the problem, splits direct from indirect injection, and gives the layered defence framework — input, tools, output, and permission plus audit.
Key takeaways
- Prompt injection is a channel problem, not a wording problem. Trusted instructions and untrusted content share one context window, and the model has no provenance-based way to tell them apart — the SQL-injection failure, one interpreter removed.
- The jump from chat to agent is what raises the stakes. Private-data access, untrusted content and an outbound path together form Simon Willison's lethal trifecta, and tool-using agents assemble it by default.
- Indirect injection is the production variant. The payload arrives inside a fetched page, a document, an email or a tool result, so a filter at the chat edge never sees it.
- Defence is layered by construction. Input screening, tool least privilege, output sanitisation, and a permission-and-audit layer that records what actually ran.
- Start from what the agent can do, not what it can say. List every tool it can call and every destination it can write to, and constrain that list before adding another detector.
What prompt injection is, and what it is not
Prompt injection is usually introduced with a definition, but the mechanism is more useful, because the mechanism is the whole problem. An application built on a language model assembles one context window from several sources: a system prompt the developer wrote, a user's message, retrieved documents or fetched pages, and tool results. The model reads all of it as one stream of tokens. There is no privileged channel for the developer's instructions; instructions and data share a format, and the model decides what to obey by inference rather than by provenance.
Prompt injection is what happens when text from an untrusted source enters that stream and is treated as though it came from the trusted one. OWASP states it plainly for LLM01: a prompt injection vulnerability "occurs when user prompts alter the LLM's behavior or output in unintended ways", and the input need not be human-visible, only parsed by the model (OWASP LLM01). The name borrows from SQL injection, and the parallel is exact: both are failures of concatenation, where intended instructions and attacker-controlled input are joined into one string the interpreter can no longer separate. Simon Willison, who coined the term in 2022, traces it to that same root cause (prompt injection series).
The distinction that trips people up is prompt injection versus jailbreaking. A jailbreak tries to make the model disregard its own safety policy; the adversary and the user are the same person, and the damage is usually to the vendor's rules. OWASP calls jailbreaking "a form of prompt injection where the attacker provides inputs that cause the model to disregard its safety protocols entirely". Prompt injection here is broader: the adversary may never talk to your app, and the target is the application — its data, its tools, its authority. jailbreaking the model's own policy takes the narrower case; this page stays with the application's surface.
Why an agent turns prompt injection into a real attack surface
In a chat-only application a successful injection usually produces a bad answer: the model says something it should not and a human reads it. The moment the same model can call tools, the blast radius changes shape, because its output is no longer only text — it is also a decision about which tool to invoke with which arguments. An injected instruction that steers that decision makes the application act under the user's own authority, often without the user seeing the step that produced it.
Simon Willison's framing makes the condition concrete. The lethal trifecta for an AI agent is (1) access to private data, (2) exposure to untrusted content, and (3) the ability to communicate externally in a way that could steal that data; an agent combining all three "can easily [be] tricked into accessing your private data and sending it to that attacker", and the reliable fix is to cut off one leg (The lethal trifecta for AI agents). The framing turns a vague worry into an inventory: name the data, the untrusted channels and the outbound paths, and see whether they meet.
Tool protocols make that inventory concrete. The Model Context Protocol gives a model a discovery call that lists a server's tools and an invocation call that runs one (MCP server tools); each tool is a capability, and a few together usually assemble all three legs without anyone intending to. OWASP files the unchecked version as its own category, "Excessive Agency — granting LLMs unchecked autonomy to take action", beside prompt injection at the top of the same list (OWASP Top 10 for LLM Applications). Injection is how the instruction arrives; excessive agency is why acting on it can hurt.
Indirect prompt injection: when the payload is fetched, not typed
The direct case — a user types a hostile instruction into your chat box — is the one everyone demonstrates, and the less dangerous one, because your application already assumes its users may be adversarial. Indirect prompt injection removes that assumption. The attacker never talks to your application; they place the payload where it will later be read: a public page the agent summarises, a document in a shared drive, an inbound email, a calendar invite, a code comment, a tool's return value. When the application fetches that content, it brings the instruction in through the front door, already trusted because it arrived as data.
OWASP defines the split in one line: "indirect prompt injections occur when an LLM accepts input from external sources, such as websites or files", and that content "may ... alter the behavior of the model in unintended or unexpected ways". The original systematic treatment is the 2023 paper by Greshake and co-authors, which showed injected instructions in retrieved content driving a model-integrated application to exfiltrate data or act without the user's request (Not what you've signed up for). It is worth reading for one reason above the rest: the attack does not require a broken model. A capable model reading a poisoned document behaves as designed and still does the attacker's bidding.
That is why the defence must include the content path, not just the chat edge. A filter inspecting the user's message sees nothing, because the message is innocent; the instruction arrives later, from a source fetched on the user's behalf, and by then it shares a context window with the system prompt. indirect prompt injection takes the variant apart — carriers, exfiltration channels and content-side controls; this centre page treats it as the fact that changes everything else.
Where the attack surface sits: the layers an agent request crosses
Prompt injection is easier to defend as several surfaces a single request crosses than as one bug. An injected instruction must enter somewhere, survive context assembly, reach a tool or an output, and produce an effect. Each is a layer a control can sit on.
| Layer | What lives there | Who controls the content | What an injection here can reach |
|---|---|---|---|
| Input | The user's message and attached files | The user, possibly hostile | The model's immediate instruction slot |
| Retrieved content | Fetched pages, documents, search and RAG passages | A third party, indirectly | Whatever the model does next |
| Tool results | Return values from the agent's tool calls | A third party or an upstream system | The next planning step and the tool it picks |
| Output and actions | Text shown to the user; tool calls with side effects | Your application | Data out, or state changed |
Two things follow. The layers a careful developer usually guards — the input layer — are the easy ones, while the layers that carry the payload in production, retrieved content and tool results, were added later when the agent got useful. And the output layer is where an injection becomes damage: text alone rarely hurts, but a tool call that writes, sends, pays or deletes is a real effect. Google Cloud's security guide names the one-layer mistake the "firewall fallacy" — inspecting the chat entrance while documents and tool output go unscanned (Mitigating indirect prompt injection in AI agents).
Prompt injection examples that generalise
The demonstrations are memorable because none is exotic; each is a template. OWASP's LLM01 entry lists two worth internalising. In the first, an attacker injects a prompt into a customer-support chatbot telling it to ignore previous guidelines, query private data stores and send email — a direct injection that escalates to unauthorised access. In the second, a user asks a model to summarise a page containing hidden instructions that make it insert an image pointing at an attacker's URL, exfiltrating the conversation through the image request (OWASP LLM01).
Several others are canonical because they were reported as real incidents: a CV carrying white-on-white instructions that a recruiting tool follows; a public issue in a source repository whose text is read by a coding agent that also holds private repo access and can open a pull request — Simon Willison notes that the official GitHub MCP server supplies all three legs of the trifecta at once (The lethal trifecta for AI agents); and, in June 2025, Google's EchoLeak, a zero-click exfiltration in which a rendered external image URL carried stolen data out of the context, fixed by a markdown sanitiser that refuses to render external image URLs (Mitigating prompt injection attacks).
The common shape is what a defence must reproduce. In every case the malicious instruction is data that was legitimately fetched: a document the user asked to summarise, an email the user received, an issue the agent was told to read. None of it broke authentication. A defence that asks only "did anything unauthorised enter" answers "no" every time, right up to the moment data leaves.
Prompt injection protection: the four defence layers
Once the problem is a stack, the defence is a stack too. The first, weakest layer is input-side screening: content classifiers that read incoming text and flag likely injections before the main model does. Google describes classifiers that detect malicious instructions "within various formats, such as emails and files", drawing on a catalogue of real adversarial examples, and Anthropic describes classifiers that "scan all untrusted content that enters the model's context window", flagging hidden text, manipulated images and deceptive UI (Mitigating prompt injection attacks, Mitigating the risk of prompt injections in browser use). Microsoft ships a comparable classifier as a service (Azure AI Content Safety). Screening helps and is not sufficient: a classifier is a detector, every detector has a false-negative rate, and an attacker who iterates will eventually produce text it does not recognise.
The second layer does not depend on recognising the payload. Tool-side least privilege decides what the agent may do regardless of what it is told, removing capabilities it does not need. If the summarising agent has no tool that writes outward, an injected instruction in a document it reads has nowhere to send anything — Willison's "cut one leg of the trifecta" turned into configuration. The guardrail products sibling surveys the screening layer; the sandboxing sibling covers the containment that runs untrusted tool code without trusting it.
The third layer is output-side handling, because injections often become exfiltration through whatever the model emits: a markdown image whose URL carries data, a link the user is nudged to click. Google's remedy for the EchoLeak class is a sanitiser, not a smarter classifier; OWASP's prevention sheet groups the same advice under escaping, output validation and constrained rendering — treat model output as untrusted to the next consumer, because it is (LLM Prompt Injection Prevention Cheat Sheet).
The fourth layer is the one teams build last and need most: permission and audit, outside the model. It decides which actions require human confirmation, which may run unattended, and what record is written when they do. Google puts a confirmation step in front of risky operations and recommends a deterministic boundary between read and mutating write actions, with an immutable audit log for every step (Mitigating prompt injection attacks). This layer is load-bearing because it never tries to tell an injected instruction from a legitimate one; it assumes both reach the model and constrains the consequence — the only claim that survives an attacker who studies your classifier. Where simulated adversaries test that claim, red-teaming a deployed agent is the sibling that covers the method.
What layered protection can and cannot promise
It is worth stating the ceiling explicitly, because the marketing around guardrails usually is not. No combination of controls makes prompt injection impossible in a system that gives one model untrusted input and consequential capability; the layers reduce the probability of success and, more importantly, bound the consequence when a detector misses. That is why this is defence in depth rather than a fix: the classifier, the least-privilege tool set, the output sanitiser and the permission-and-audit layer each fail differently, and a payload must survive all of them to cause harm.
Two honest consequences follow. The guarantee you can actually make is a containment promise — "no injection can cause an action we did not permit" — which is far stronger than "no injection will occur" and the only one worth writing down. And the usefulness-versus-safety trade-off is real: removing every outbound capability removes the exfiltration risk and much of what the agent was for. The design work is deciding, per task, which capabilities are worth the risk and which need a human in the loop. The NIST AI Risk Management Framework is useful here less for its controls than for its discipline, organising the work into Govern, Map, Measure and Manage so the risk decision is deliberate rather than defaulted (NIST AI RMF).
Prompt injection detection: signal, not proof
Detection deserves its own treatment because it is the control people reach for first and over-trust. Detection produces a signal — "this input looks like an injection" — and a signal is not proof in either direction. A flag can be a false positive that breaks legitimate content, and a miss is a false negative that ships the payload. The current generation is trained and imperfect: Google notes its classifiers draw on a curated catalogue of real-world adversarial data, and Anthropic is candid that its own browser agent is not immune, publishing an attack-success-rate chart to show progress "not to claim the problem is solved" (Mitigating the risk of prompt injections in browser use).
The practical place for detection is one input to a decision, never the decision. NIST's Generative AI Profile files prompt injection under its information-security risk category, beside model theft and unsafe tool use, and maps responses to the same Measure and Manage functions as the rest of the framework — measure the rate, manage the consequence, and do not mistake the measurement for a barrier (NIST AI 600-1). MITRE ATLAS is a useful companion for naming the adversary techniques a rule tries to match, so coverage can be discussed as coverage (MITRE ATLAS). Where those signals become concrete upstream checks is the cluster's OWASP Top 10 walkthrough.
LLM prompt injection at the model layer
A fair question is what the model vendors are doing, since a reader may hope the model itself will be hardened enough to remove the problem. Meaningful work is shipping here and it reduces risk without eliminating it. Google describes a layered approach "introducing security measures designed for each stage of the prompt lifecycle", from model hardening through classifiers to system-level safeguards, and reports that adversarial training measurably improved its flagship model's resistance to indirect injection; Anthropic describes reinforcement-learning training that rewards the model for refusing injections in simulated web content, plus classifiers that adjust behaviour when an attack is detected (Mitigating prompt injection attacks, Mitigating the risk of prompt injections in browser use).
The load-bearing word is "stage". Provider hardening raises the cost and catches a large share; it does not know which documents your application fetches, which tools it exposes or which data it can reach. Anthropic's builder guidance makes the split practical: deliver third-party content inside tool-result blocks rather than the system prompt, tell the model what the content is and where it came from, and state that tool-returned content is untrusted data (Mitigate jailbreaks and prompt injections). That is a description of the application's job, and it is the application's job precisely because the model layer cannot do it from the outside.
Where SmartGate fits
SmartGate does not detect or block prompt injection, and this page will not pretend otherwise. What it provides is one of the four layers above — the permission-and-audit layer, and the metering that makes an agent's behaviour accountable afterwards. SmartGate is an MCP-native algorithm gateway for token control, traffic shaping and agent audit: a single authenticated endpoint through which an agent reaches its tools, with per-key metering and an audit record written as calls happen. Seven tools are exposed through it — smart_fetch, smart_search, smart_context_gate, smart_dedup, smart_budget_guard, smart_memory and smart_pipe — and each call is counted against the caller's key.
That matters to prompt injection for a specific and modest reason. An injection that makes an agent loop, over-fetch or call a tool thousands of times is a cost and containment problem before it is an exfiltration problem, and the control plane is where that becomes visible and capped rather than discovered on an invoice. Traffic shaping and per-key limits bound the blast radius of a runaway agent whatever the reason, and the audit record lets a team reconstruct what was called and by whose key — the Manage half of the incident, not the prevention half. The surfaces are documented rather than described here: call control on token control, the record on audit and compliance, and the connection steps in connect an MCP client. Plan limits move with the tier — monthly token caps of 2M, 20M, 100M and 200M+, MCP requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days — so the pricing page is the authoritative table.
The prompt injection surface, read from the source
Every layer above is somebody's published work. This one is ours, read at origin/main on 2026-10-08,
and it answers a narrower question than whether an injection is stopped — SmartGate does not detect or
block prompt injection — namely: where the injection surface sits in the architecture, and what shape
external content takes before the model reads it.
One path in, and it is stateless. backend/smartgate/api/mcp.py builds the server as an MCP
FastMCP instance with stateless_http=True, json_response=True and streamable_http_path="/", so
the MCP endpoint is a single streaming-HTTP path reached over the authenticated transport (MCP
2025-03-26) rather than a set of ad-hoc routes. The stateless part matters for this page's threat
model: there is no server-side session carrying context between calls, so what a call can act on is
what the call itself presents, counted against the caller's key.
One audit hook per tool call. Every tool handler returns through _run_with_audit(tool, ctx, process_coro, params), which first calls _ensure_mcp_audit_context() — binding route="mcp" and an
event kind of tool, with agent_platform defaulting to the request's own when it is present — then
awaits the app's audit_hook exactly once, after the handler and before the result is handed back as
JSON. A failed tool raises ToolError from the same place. That is the permission-and-audit layer of
the four above, located in the source: the call and its record are written at one point, and a tool
cannot finish without passing through it.
The text the model reads about a tool has one origin. The tool descriptions and annotations are not
written at the handler; they are imported from smartgate.api.mcp_tool_docs as TOOL_DESCRIPTIONS and
tool_annotations, and each registration passes them in — description=TOOL_DESCRIPTIONS["smart_fetch"]
and annotations=tool_annotations("smart_fetch"), and the same for the other six tools. For a page
about instructions arriving through a channel the developer does not control, this matters at the seam:
the metadata the model reads is a single reviewed source rather than text scattered across handlers, so
what a model is told about a tool can be diffed like any other file.
The shape of external content before it enters context. backend/smartgate/modules/fetch/html_converter.py
turns a fetched page into Markdown through SmartHTMLConverter.convert(html): markdownify with ATX
headings, - bullets and script, style, nav, footer and aside stripped, then a GFMProcessor
pass that closes code fences, inserts table separator rows and normalises task lists. The readout is a
description of form, not a defence: the external text arrives in the context as a normalised Markdown
string, and the boundary this page draws between trusted instructions and untrusted content is not a
filter at this step.
What this does not claim. None of this inspects a payload. The single endpoint, the per-call audit hook, the single metadata source and the Markdown shape are the places a control could sit and the places our record of what ran does sit; whether an injected instruction survives them is the four-layer question above, and this page's answer to it stays containment, not detection.
How to get started
The useful first move is not buying a detector; it is drawing the surface, because the surface is what any control sits on.
- List the untrusted channels. Every source of text the agent reads that a third party can influence: fetched pages, uploaded documents, inbound email, issue trackers, tool returns. That list is the injection surface.
- List the private data and the outbound paths. The data the agent can reach and the ways it can send anything out — HTTP, email, a rendered link or image. If both are non-empty, you have the lethal trifecta and should assume it will be exercised.
- Cut the easiest leg first. Reduce outbound capability and split read from write, so an injected instruction has nowhere to send what it gathers. This is configuration, not procurement, and it is the highest-value step.
- Put a confirmation step in front of mutating actions. Anything that writes, pays, deletes or sends should require explicit approval, and the approval should be logged.
- Then add detection and guardrails as one input. Read the guardrail-products sibling for where they sit, and wire the audit trail so you can tell whether the rates move.
- If you route tool calls through a gateway, start free and watch one call end to end before trusting the limits; the pricing page tells you which tier real volume needs.
Frequently Asked Questions
Is prompt injection the same as jailbreaking?
No, though the terms are often used interchangeably. Jailbreaking is a user trying to make a model disregard its own safety policy, so the adversary and the user are the same person and the harm is usually to the vendor's rules. Prompt injection is broader: the adversary may be a third party who never talks to your app, and the target is your application, its data and its tools. Jailbreaking is a subset of the wider problem.
Can we just filter malicious instructions at the input?
Not with input filtering alone. It is the first and weakest layer: it can catch known patterns and it will miss a determined attacker, because a detector is not a barrier. The control that does not depend on recognising the payload is least privilege at the tool layer, plus a confirmation step in front of actions with side effects.
How is indirect prompt injection different from a direct one?
In a direct injection the attacker types the hostile instruction into your system, so a filter at the chat edge can at least see the attempt. In an indirect injection the instruction is planted where your application will later fetch it, such as a web page, document, email or tool result, and it arrives already treated as trusted data. Nothing the user typed is malicious, which is why indirect attacks slip past input checks.
What is the single most effective control?
Cutting one leg of the trifecta, usually the outbound one. If the agent cannot send data out, an injected instruction has no exfiltration path, and that holds regardless of whether any detector notices the payload. It is configuration rather than a purchase, and it is the first thing to do.
Where does detection fit if it is not a barrier?
Use it as one input to a decision, never as the decision. Detection tells you a rate is moving and helps you triage, and it works best backed by containment that limits the consequence of a miss.
Limitations
This page is a framework, not a benchmark, and it names no guardrail product's performance. The four-layer model is a classification aid: real systems overlap the layers, a single component sometimes screens input and output at once, and the right place for a control depends on where your context is assembled.
Layered defence reduces risk and never reaches zero in a system that combines untrusted input with consequential capability; the achievable claim is containment, not immunity. And this page makes no claim about any specific product's detection or blocking behaviour, including ours — SmartGate's role is described narrowly and only where this project records the facts.
The external descriptions above are each source's own published wording, read at the linked pages, and they describe scope rather than quality. Where a source names a technique without a published measurement, this page supplies no number. The demand figures are this project's own measurement.
Sources
- OWASP — Prompt Injection (LLM01) and the Top 10 for LLM Applications index.
- OWASP — LLM Prompt Injection Prevention Cheat Sheet.
- Greshake et al., 2023 — Not what you've signed up for.
- Google — Mitigating indirect prompt injection in AI agents and Mitigating prompt injection attacks with a layered defense strategy.
- Anthropic — Mitigating the risk of prompt injections in browser use and Mitigate jailbreaks and prompt injections.
- Microsoft — Prompt injection attacks and mitigations and Azure AI Content Safety jailbreak detection.
- Simon Willison — The lethal trifecta for AI agents and the prompt injection series.
- NIST — AI Risk Management Framework and the Generative AI Profile (NIST AI 600-1).
- MITRE — ATLAS.
- Model Context Protocol — server tools.
- The demand figures are this project's own paid measurement: DataForSEO Google Ads, United States,
12-month window, measured 2026-10-04, recorded in this project's
search_volume.jsonandresearch_brief.md.
Method note
This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned none of this page's seven sections (0 abstentions, 7 no-slice verdicts): rule A found no unique symbol in the scanned repository for any section keyword, and the remote candidate fallback returned generic helpers that are collisions rather than section-specific evidence. A pinned generic would have given the page the shape of a verified article with none of the substance, so every section is written from external, linkable sources, which is the house rule for an unpinned section.
Product facts were read read-only from the product source at the revision the slice run recorded in
this project's pipeline_results.json, and the plan figures were re-checked against the live pricing
page on 2026-10-04; the demand figures are this project's own measurement. Every external statement
quoted above is taken from the URL cited beside it. No code, batch fingerprints, auction data or
internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.
The slice run for this page recorded 0 of 7 sections pinned, 0 abstention(s) and 7 no-slice verdict(s); BLOCKS is empty because the matcher found no unique symbol for any section rather than section-specific evidence, as the Method note above explains.