SmartGateSmartGate

Indirect Prompt Injection: the Path from Content to Action

Indirect prompt injection happens when an agent reads content a third party controls — a fetched page, an email, a document or a tool result — and treats part of it as an instruction, so its next action serves the attacker. No malicious user is involved and no model is broken: the text shares one context window with the system prompt, and the model cannot verify where any of it came from.

Short answer: Indirect prompt injection happens when an agent reads content a third party controls — a fetched page, an email, a document or a tool result — and treats part of it as an instruction, so its next action serves the attacker. No malicious user is involved and no model is broken: the text shares one context window with the system prompt, and the model cannot verify where any of it came from. The control that does not depend on spotting the payload is limiting what the agent is allowed to do.

Key takeaways

  • Indirect injection is a path, not a prompt. The payload travels from a carrier the agent reads to an action the agent takes; every control that works sits on one of the three hops.
  • The data/instruction boundary is behavioural, not structural. Labels and "treat this as data" instructions change the odds without creating a channel the model can trust.
  • The model cannot self-verify provenance. Origin is metadata that lives in the application, not in the text, so it is gone before the model sees the string.
  • Least privilege is the primary defence. It holds when detection misses, because a capability the agent does not hold cannot be invoked by any instruction.
  • Start by drawing the path for one running agent: list its carriers, the context they enter, and the actions it can reach, then cut the capability the payload needs.

The propagation path: content, context, action

An indirect injection is easier to defend once it is drawn as a path. The path has three hops, and every control that works sits on one of them.

  1. Carrier to context. Content a third party can influence is fetched, received or returned, and the application assembles it into the model's context window beside the system prompt and the user's message.
  2. Context to decision. The model reads the whole window as one stream of tokens and chooses the next step — an answer, or a tool call with arguments.
  3. Decision to action. That step runs: a message is sent, a record is written, a payment moves, a link or a rendered image leaves the browser.

The hop people picture is the second, where the text is "understood". The hop that does the damage is the third, where an instruction becomes an effect. The hop hardest to guard is the first, because the content is supposed to be there: the agent fetched the page, received the email, or ran the query. Direct injection starts at a door already treated as adversarial — the user's message — which is why the indirect case needs a different mental model, not a stricter filter on the same door.

This page asks, at each hop, which control changes the outcome, while the centre page covers the problem itself and the four-layer defence framework — the definition and the layered framework; here the focus is the content that travels and the authority it borrows when it arrives.

Where untrusted content enters an agent's context

The first hop has the most doors, and they have multiplied as agents became useful. A chat assistant reads what a user types; a tool-using agent reads everything its tools bring back. None of the following requires the attacker to hold an account on your system:

Carrier Who can influence it How it reaches the window
Fetched web page The page's author, or anyone who can edit it A fetch or browse tool
Inbound email Any sender A mail tool the agent reads
Calendar invite or meeting note Any invitee A calendar tool
Uploaded or shared document The document's author A file, drive or upload tool
Search and retrieval passages Whoever ranks in the index A search or RAG step
A tool's returned value The upstream system, or a third party behind it The result of a tool call
Issue, ticket or code comment Any contributor A tracker or repository tool
Image alt text and metadata The image's host A vision or render step

Two entries are trusted by default and deserve emphasis. A tool's return value looks like the agent's own machinery reporting back, so an instruction inside it reads as a system message rather than as third-party content. Retrieval is chosen by similarity rather than trust, so a poisoned passage that is merely relevant is pulled into the window without anyone deciding it was safe.

The entry surface is therefore the union of every tool the agent holds: the fetch tool that makes summarisation useful is also a delivery mechanism for a page written to talk to summarisers. You cannot shorten the carrier list without removing capability, so the useful question is not how to stop content entering but what the agent may do once it has.

Data as instruction: where the boundary actually sits

A common hope is that the boundary between data and instruction is structural — that a delimiter, a label or a tool-result wrapper tells the model which text is which. It is not. A language model reads one token stream, and the developer's instructions and the attacker's payload arrive in the same format. OWASP states the point in its LLM01 entry: an injection need not be human-visible, only parsed by the model, because the model cannot tell an instruction from the text of an instruction unless something outside the model says so.

Wrappers and delimiters are still worth using, and their value is bounded. Telling the model "the following is untrusted content" moves the odds, because a capable model often complies; it does not create a channel the model can trust, because the same window can contain a later sentence that withdraws the warning. Research on the instruction hierarchy tried to make that boundary a property of training rather than of phrasing, and it improves behaviour without authenticating who sent anything — training shapes priorities, it cannot verify provenance.

Work on securing agents against injections reaches the same conclusion from the design side: the patterns that keep an agent from acting on a compromised observation are architectural — the action is chosen from a fixed set, or the agent is denied the capability the injected instruction needs — not prompt-phrasing tricks (Beurer-Kellner and co-authors, arXiv:2506.08837). The boundary that holds sits outside the model, in the code that assembles the window: only that code knows which bytes came from the system prompt and which from a page a stranger wrote.

Why the model cannot verify trust from inside the context

It is worth being precise about why the model cannot simply "check the source". Provenance is not a property of the words. The sentence "Ignore your previous instructions" is identical whether a user typed it, a developer wrote it into a test fixture, or a web page buried it in white text; only the metadata about where it came from differs, and metadata lives in the application, not in the tokens the model reads. Asking the model to distrust untrusted text is asking it to recover information removed before it saw the string, which is why every prompt-level instruction to "ignore anything suspicious" is a probability, not a guarantee.

This also separates indirect injection from the problem next door. A jailbreak is a user trying to make a model disregard its own rules: the adversary and the user are the same person, and the harm is usually to the vendor's policy. An indirect injection has a different adversary — a third party who never talks to the application — and a different target: the application's data, tools and authority. The distinction matters because the controls differ: turning a model against its own policy is largely a model-hardening problem, while an injection through fetched content is an application-architecture problem whose fix lives in the plumbing around the model.

A team that builds its defence around the model noticing a bad instruction is building on a capability the model does not have, and a detector at the chat edge checks the one hop where the indirect route does not enter. A team that bounds what any instruction can reach works from the facts known at assembly time: which capabilities the agent was handed and which destinations it can write to.

From context to a tool call: the hop that turns text into action

In a chat-only application the second and third hops collapse into one: the model answers and a human reads. When the model can call tools, the output is a decision — which tool, with which arguments — and an injected instruction that steers that decision borrows the user's own authority. The Model Context Protocol makes the shape explicit: a client discovers a server's tools with a listing call and runs one with an invocation call, so a tool is a capability the agent holds and an injection is an attempt to choose it (MCP server tools). The model is not "obeying the web page"; it is choosing, from its tool list, the action the page's text has made look reasonable.

Simon Willison's lethal trifecta names the condition under which this becomes exfiltration. An agent that combines access to private data, exposure to untrusted content, and a way to send data out can be steered into doing exactly that, and the reliable fix is to remove one leg rather than to detect the instruction. A tool-using agent usually assembles all three by accident: the fetch tool supplies the untrusted content, the filesystem or database tool supplies the private data, and a mail or HTTP tool supplies the way out. None is a vulnerability on its own; together they are a route, and the route is what the payload wants.

The inventory doubles as a checklist: name the data the agent can read, the carriers it reads, and the exits it can use, then remove the leg that costs the least capability — usually the outbound one.

Indirect prompt injection examples that show the path

Useful examples show the hops, not a clever string.

The BIPIA benchmark measured the indirect case directly, injecting instructions into the external content a model reads — an email it is asked to summarise, a web page it reasons over — and found the attacks succeeded across tasks, the empirical version of "the content is the attack surface" (Yi and co-authors, arXiv:2312.14197). InjecAgent aimed at tool-integrated agents, showing that a poisoned observation returned by one tool can redirect the next tool call a user never requested (Zhan and co-authors, arXiv:2403.02691). Both share a structure: nothing the user typed was hostile, and the model behaved exactly as designed.

The original systematic treatment remains the clearest on mechanism. Greshake and co-authors showed that injected instructions in retrieved content could drive a model-integrated application to act without the user's request, and their point was that the attack needs no broken model — a capable model reading a poisoned document does the attacker's bidding while working correctly (arXiv:2302.12173). The June 2025 EchoLeak class of issues showed the final hop in its purest form: an instruction that made the assistant render an external image URL, which carried stolen data out in the request itself, fixed by refusing to render external image URLs rather than by improving the classifier (Google security blog).

Read together these are a map of the three hops: legitimate content the agent was asked to read, a context window with no way to mark it foreign, and an action the application had permitted. Every one would have been blunted by the same control — not giving the agent the capability the payload needed.

llm prompt injection: the same path at the protocol and model layer

The path does not disappear when the agent reaches models and tools through a protocol; it moves, in two ways.

First, tool descriptions and tool results become part of the surface. A tool's description is text the model reads to choose a tool, and a tool's result is text the model reads to decide the next step; both can carry instructions, and the result in particular arrives looking like first-party output. The protocol's own security guidance treats a server as untrusted until shown otherwise: the tool says what it says, and the application decides what to believe. The builder-side counterpart is to deliver third-party content in tool-result blocks rather than in the system prompt, and to state that tool-returned content is untrusted data — an instruction addressed to the application, because the model cannot make that determination from the inside (Anthropic builder guidance).

Second, model vendors have moved part of the problem into the model. Google describes hardening, input classifiers and system-level safeguards across the prompt lifecycle, and reports that adversarial training measurably improved resistance to indirect injection; Anthropic describes training that rewards a model for refusing injections planted in simulated web content (Google security blog; Anthropic research). Provider work raises the cost without knowing which pages and tools your application uses. Model-layer hardening reduces how often a steering succeeds; it does not bound the consequence when one does, because that is a property of the tools you attached.

Prompt injection prevention: least privilege as the primary defence

Prevention, for the indirect path, is mostly not the word the market uses. It is not a filter that inspects content on the way in, because the content on the way in is supposed to be there. It is a decision about what the agent is permitted to do, made once, in configuration, and enforced outside the prompt.

Least privilege is the primary defence for a specific reason: it does not depend on recognising the payload. A classifier is a detector, every detector has a false-negative rate, and an attacker who iterates eventually produces text the detector has not seen. A capability the agent does not hold cannot be invoked by any instruction, however well written, and that is the property that survives an adversary who studies your defences. The strongest single move is therefore to remove a capability the task does not need rather than to add a second detector in front of the one it does.

Three constraints do most of the work. Read and write are split, so an agent that reads untrusted content cannot also mutate state. The outbound path is narrowed, so an agent that must write outward writes to one permitted destination rather than to any address a payload suggests. And duty is scoped, so the credential an agent holds reaches one service rather than everything — scoping the credentials a single agent holds is the identity half of the same idea.

The layer above the tools is confirmation. Anything with a side effect that cannot be undone — a payment, a deletion, a message to a third party — should require an explicit human step, and the step should be logged. The friction is the point: it turns an unattended injection into a decision someone made. Where the untrusted code a tool runs must be executed at all, isolating it is a separate control — running tool code inside an isolation boundary limits what a compromise can touch, independent of whether any detector fired. Screening still has a place, as one input to a decision and as a rate to watch; it is covered on its own terms by the screening products and where they sit, and should never be the only layer.

Cutting the path at each hop: a control map

Hop A control that fits What it removes Residual risk
Carrier to context Fetch allowlists; treat tool returns as untrusted Some delivery routes A page the task legitimately needs still enters
Context to decision Provenance labels; third-party content in tool-result blocks Some ambiguity about what is data The model can still be argued out of a label
Decision to action Least-privilege tools; read/write split The capability the payload wants An allowed capability can still be abused
Action to effect Confirmation for irreversible acts; scoped credentials Unattended damaging writes Friction, and a fatigue risk in the human loop
Any hop Screening and detection Some known payloads False negatives, by construction

The left column is where an attack lives; the middle column is what can be done there; the right column is what is left over, and no row has an empty right column. Defence in depth is not a slogan here — each control fails differently, so a payload must defeat every layer between it and an effect. The ordering that matters is to prefer the control that does not depend on recognising the payload: least privilege and confirmation come before classifiers, and screening the user's message while documents and tool results go unseen is a guard placed where the indirect payload never lands.

A prompt injection cheat sheet for the content path

A short list to check against a real agent, written for the indirect path:

  • Enumerate the carriers. Every tool that returns text is a carrier; write the list down and treat it as the injection surface.
  • Mark provenance in the assembler, not the prompt. The application knows where each block came from; make it say so, and keep third-party content out of the system prompt.
  • Widen the blast radius on purpose. List every tool the agent can call and every destination it can write to, then delete the ones the task does not need.
  • Split read from write. An agent that reads untrusted content should not hold a tool that mutates state unattended.
  • Remove the outbound leg. If nothing can be sent out, nothing can be exfiltrated, whatever an instruction achieves inside the window.
  • Confirm the irreversible. Payment, deletion and messages to third parties deserve an explicit, logged human step.
  • Treat detection as a rate to watch, not a wall to hide behind.
  • Re-check after every capability you add — the threat model changes when a new tool is granted.

The OWASP prompt injection prevention cheat sheet is a useful companion for the controls at the input and output ends of this path — escaping, validation and constrained rendering. (OWASP prevention cheat sheet)

Where SmartGate fits

SmartGate does not detect or block prompt injection, and this page will not claim it does. What it provides is the layer that bounds the consequence: a single authenticated endpoint through which an agent reaches its tools, with per-key metering and an audit record written as calls happen. SmartGate is an MCP-native algorithm gateway for token control, traffic shaping and agent audit, and three of its behaviours matter here.

Its per-key rate limits and per-team budget caps bound the cost of a runaway loop, the first thing an injected instruction usually causes: an agent told to fetch, search or call a tool in a loop bills before it exfiltrates. Its traffic shaping narrows what a caller can do at the edge. And its audit record lets a team reconstruct what was called, by which key, and when — the Manage half of an incident rather than the prevention half, covered in more depth by the operator's key, log and quota view.

The plan limits move with the tier — monthly token caps of 2M, 20M, 100M and 200M+, MCP requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and 2, 10, 30 or unlimited team keys — so the pricing page is the authoritative table. Call control and the record are documented on token control and audit and compliance, and neither substitutes for the least-privilege work above.

Prompt injection prevention: what the fetch path keeps and what it strips

When the gateway turns a URL into Markdown, the defence that matters here is not in the prompt; it is in what the converter physically removes before any text can reach a model context. SmartHTMLConverter.convert in backend/smartgate/modules/fetch/html_converter.py calls markdownify with a single stripping list — strip=["script", "style", "nav", "footer", "aside"] (line 91) — so those five element types never appear in the output. Everything else survives by default: the same call keeps links and images, which is why a payload riding in an anchor target, an image alt string or an inline attribute is carried through intact. The GFM pass that follows (GFMProcessor.process, line 63) only closes code fences, inserts missing table separator rows and normalises task-list markers; it makes no security judgement.

The surrounding pipeline in backend/smartgate/modules/fetch/algorithm.py shapes the rest of the context. _clean_html (line 190) tries trafilatura.extract with include_formatting, include_links and include_images all true (lines 194–200), and its lxml fallback deletes script, style, nav, footer, aside and, on top of those, header (line 206). Metadata is extracted from the raw HTML before cleaning (_extract_metadata, lines 211 and 298) and returned beside the Markdown, so OpenGraph tags, the description, JSON-LD and a favicon are part of the fetch result even though the body was sanitised. Two hard bounds sit in front of all of it: _validate_url (line 63) rejects non-HTTPS schemes and any host that resolves into a private or loopback range, and MAX_RESPONSE_BYTES caps a response at two megabytes (line 21). The honest reading is that script and styling are gone, but link targets, alt text and page metadata are exactly the carriers that remain — which is why least privilege, not conversion hygiene, is the control that bounds the damage.

How to get started

The first move is not to buy a detector; it is to draw the path for one agent that is already running, in an order where each step makes the next cheaper to reason about.

  1. Pick one agent and list its carriers. Every tool that returns text: fetch, search, mail, files, tickets, code. That list is the surface.
  2. List the data and the exits. What private data the agent can reach, and every way it can send something out. If both are non-empty, assume the route will be exercised.
  3. Cut the capability the payload needs. Turn off the outbound tool, split the write tool from the read tools, and scope the credential to the one service the task uses. This is configuration, not procurement, and it is the highest-value step.
  4. Add confirmation to irreversible actions. Payment, deletion and third-party messages get a human step and a log line.
  5. Only then add screening and detection, as one input to a decision and as a rate to watch.
  6. Watch the call record. Start free and read one call end to end before trusting the limits; the pricing page tells you which tier real usage needs.

Frequently Asked Questions

What makes a prompt injection "indirect"?

An indirect injection is planted where the application will later read it rather than typed into the chat. The attacker is a third party who never talks to your system: they put the payload in a page, an email, a document or a tool's output, and your agent fetches it on the user's behalf. Nothing the user typed is hostile, which is why a filter at the chat edge does not see it.

If a model is hardened against injections, is the problem solved?

No. Provider hardening raises the cost of a successful steering and does not know which pages your application fetches or which tools it exposes. Hardening reduces how often an instruction is followed; it does not bound what happens when one is followed.

Why is least privilege better than a content filter here?

Because it does not depend on recognising the payload. A filter is a detector, detectors have false negatives, and an attacker who keeps trying will eventually produce text yours has not seen. A capability the agent does not hold cannot be invoked by any instruction, so containment holds whether or not any detector fires.

Can we just wrap untrusted content in a tag the model is told to treat as data?

It helps and it is not a boundary. The tag changes the odds that a capable model treats the passage as data, but it is still text in the same window as the system prompt, and a determined payload can argue with it. Enforcement has to live outside the model, in what the agent is allowed to do.

Where does detection fit if it is not a barrier?

Use it as one input to a decision and as a rate to watch. Detection tells you attempts are happening and helps you triage, and it works best on top of containment that limits the consequence of a miss. It should never be the only layer.

Limitations

This page is a path model, not a benchmark, and it names no guardrail product's accuracy. The three-hop split is a classification aid: real systems overlap the hops, a single component sometimes both fetches content and acts on it, and the right place for a control depends on where your context is assembled.

Least privilege bounds consequence and does not remove the risk of an injection occurring; the achievable claim is containment — no injection can cause an action we did not permit — not immunity. Removing every outbound capability removes the exfiltration route and much of what the agent was for, so the design work is deciding, per task, which capabilities are worth keeping.

This page makes no claim about any specific product's detection or blocking behaviour, including ours — SmartGate's role is described narrowly and only where this project records the facts. The external descriptions are each source's own published wording, and they describe scope rather than quality. The demand figures are this project's own measurement.

Sources

Method note

This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned none of this page's seven sections (0 abstentions, 7 no-slice verdicts): rule A found no unique symbol in the scanned repository for any section keyword, and the remote candidate fallback returned generic helpers that are collisions rather than section-specific evidence. A pinned generic would have given the page the shape of a verified article with none of the substance, so every section is written from external, linkable sources.

Product facts were read read-only from the product source at the revision the slice run recorded in this project's pipeline_results.json, and the plan figures were re-checked against the live pricing page on 2026-10-04; the demand figures are this project's own measurement. Every external statement quoted above is taken from the URL cited beside it. No code, batch fingerprints, auction data or internal hosts are transcribed.

The slice run for this page recorded 0 of 7 sections pinned, 0 abstention(s) and 7 no-slice verdict(s); BLOCKS is empty because the matcher found no unique symbol for any section rather than section-specific evidence, as the Method note above explains.