AI Agent Security: The Runtime Surface Around Every Tool Call
AI agent security is the runtime question, not the prompt-writing one: what each tool may do and as whom, which MCP servers the agent trusts, how untrusted content is marked before it reaches the model, and where keys and audit rows live. Every control sits outside the model, which cannot separate your instructions from text it was asked to read.
Short answer: AI agent security is the runtime question, not the prompt-writing one: what each tool may do and as whom, which MCP servers the agent trusts, how untrusted content is marked before it reaches the model, and where keys and audit rows live. Every control sits outside the model, which cannot separate your instructions from text it was asked to read.
ai agent security: the runtime surface, not the input filter
An agent is not a chatbot with more features. It is a program that holds credentials, picks its own next action, and reaches systems your application already trusts. The OWASP agent cheat sheet states the consequence bluntly: an agent's blast radius equals every credential, tool and API it can touch, and one bad step compounds across an entire plan instead of staying inside a single response.
Four questions cover most of that surface, and they are the four movements of this page. Tool permission boundaries — which functions the agent may call, and under which identity. MCP server trust — which external servers are allowed in, and how their tool definitions are pinned. The injection surface — which content the agent reads, and how that content is marked as data rather than instruction. Credentials and the record — where keys live, and what an audit row can prove after an incident.
Two structural facts make this different from ordinary API security. First, a language model cannot reliably separate trusted instructions from untrusted data once both are concatenated into one context window. That is why the fix is architectural rather than a better system prompt. Second, authorization has to be enforced by something that is not the model: OWASP's Excessive Agency entry calls this complete mediation — validate every downstream action against policy, rather than asking the model to decide whether an action is allowed.
The combination to treat as dangerous is narrow and worth naming, because it tells you which agent to fix first: the agent can read private data, it can read content an outsider controls, and it has a channel that leaves your boundary. With all three present, an indirect injection needs no exploit code: the agent is the exfiltration path. Remove one leg by design and the finding drops from "incident" to "log line". Input-side hygiene for the same request path — normalising the policy, validating the key, masking before logging — belongs to secure prompt handling in AI applications.
mcp security: four attack classes the protocol names itself
You do not have to invent the MCP threat model: the protocol's own Security Best Practices document names the attack classes, and every one is a runtime problem.
Confused deputy. An MCP proxy that holds one static client identity and connects to a third party can be walked into obtaining authorization codes without the user's genuine consent, by combining a static client ID, dynamic client registration and a consent cookie. The specification's answer is a per-client consent step and a registry of approved client identifiers checked before the third-party authorization flow starts.
Token passthrough. An MCP server that accepts tokens it did not validate as issued to it, and then forwards them downstream, breaks the OAuth audience boundary and creates the classic confused deputy at the next hop. The specification forbids it: a server must not accept tokens that were not explicitly issued for it, and must validate the audience claim. The failure mode is broader than theft: rate limits, request validation and monitoring that keyed off the audience stop applying.
Session hijacking and event injection. If a server hands out a session identifier and accepts it from any caller, an attacker who obtains that identifier impersonates the original client. The mitigations are concrete: session identifiers must be cryptographically random, bound server-side to the authenticated principal (keying stored state by user and session, where the user part comes from the verified token), and must not be treated as authentication in themselves.
Local server compromise. A locally configured server is a process running with the client's privileges, and the attack surface is the configuration itself: a malicious startup command, or a server left listening on an interface it should not be. Hence the requirements on one-click configuration — show the exact command, untruncated, before executing it; require explicit approval; run the server in a sandbox with minimal privilege.
What to check: pull the rejected-request side of your logs. Authentication failures should be separated by cause — audience mismatch, unknown client identifier, missing or foreign session identifier — with at least one real attempt of each in your test environment. One generic error code for all three tells you nothing you can act on.
mcp server security: what a server has to refuse
The client side asks "can I trust this server". The server side asks a blunter question: what do I refuse, regardless of how persuasive the caller is. Three refusals carry most of the weight.
Refuse tokens that were not issued for you. This is the audience check above, seen from the server: it must not accept a token merely because the token is valid somewhere, and must not treat possession of a state handle as proof of identity. Both are required by the specification and both are routinely skipped, because a happy-path test never exercises them.
Refuse to be an open process. A server intended for local use should bind only the local interface, validate the Origin header on incoming Streamable HTTP connections, and require the Host header to be one it expects. The specification lists all three, because without them a remote page can reach a locally running server through a name that resolves differently at check time than at use time. The request is well formed and TLS-valid — it is simply from somewhere it should never have come from.
Refuse to keep a client's secrets. OWASP's guide to third-party MCP servers makes this a design rule rather than a hygiene tip: run untrusted servers in a container with bounded filesystem and network reach, grant tool permissions just-in-time and scoped to one task, and prevent the server from reading client-side tokens, history or cached memory.
Two integrity controls belong here as well, both aimed at the delivery channel rather than the request: pin each tool definition by hashing its canonical name, description and input schema and re-hash on every discovery, so a definition that mutates after approval is visible as drift rather than silently adopted; and treat the whole schema as an injection surface, not just the description field, enforcing strict parameter schemas that reject undeclared properties. OWASP names the technique under tool poisoning and the rug-pull variant, where a trusted tool is replaced by a malicious definition that inherits the original's approved reputation.
That pinned definition is also a published contract, and keeping the reference — titles, annotations and schemas — true to the code that serves it is covered on API documentation best practices for AI tools.
What to check: you want a drift alarm, not a drift report. Store the schema hashes where the approval happened, re-compute them on each connection, and fail closed on mismatch, with the server and tool name attached to the log line. Then point the server at a token minted for a different audience and confirm the refusal is recorded as an audience failure, not as a generic error.
prompt injection protection: why the boundary cannot be the prompt
Prompt injection is the vulnerability class where the caller supplies text the model then treats as instruction. It splits into two shapes that need different controls. Direct injection is the user's own input trying to override the system message. Indirect injection — the one that matters for agents — is instruction smuggled into content the agent reads on someone else's behalf: a document, an email, a web page, a tool response. Anthropic's browser-agent research puts it plainly: no agent that processes untrusted content is immune, and every page and embedded document is a potential vector.
The measured defence is to mark provenance, not to filter vocabulary. Microsoft's spotlighting research transforms untrusted content so it carries a continuous signal of where it came from — delimiting, interleaving a marker through the text, or encoding it — and reports attack success falling from above 50% to below 2% on GPT-family models with negligible loss on the underlying task. Delimiter-only is the weakest of the three and is subverted by an attacker who learns the system prompt; encoding is the strongest and needs a capable model.
Productised classifiers help and are not a boundary. Microsoft's Prompt Shields separates user prompt attacks from document attacks and scans both at the user-input intervention point, with document attacks also scanned on the tool response path; each request returns a detected flag and a filtered flag. Classifier evasion is a live field, which is why Microsoft's own zero-trust guidance pairs the probabilistic layer with deterministic ones. Anthropic's mitigation guidance takes the same layered shape: a lightweight screening call constrained to a structured yes/no verdict, hardened system instructions, and explicit handling of untrusted tool content.
Three deterministic controls are what actually hold when a classifier misses. Mark untrusted content and never place it in a system-role message, because a system-role slot is the one place the model is trained to obey. Keep authorization out of the model — complete mediation, per OWASP — so an injection that succeeds in changing the plan still cannot execute an action the identity is not entitled to. And put a human on the irreversible steps: approvals, deletes, payments, outbound sends.
What to check: measure your own attack success rate instead of assuming one. Keep a small corpus of injection payloads aimed at your real tools, run it against each revision of the prompt, the tool descriptions and the retrieval path, and record which layer stopped each one — the classifier, the tool-schema validation, the permission check, or the approval step. If every payload is blocked by the same layer, you have just measured a single point of failure.
mcp security best practices, in the order they pay off
The controls above are cheaper in one sequence than in another: each step makes the next one cheaper, because it produces the artifact the next step reads. This is the order that survives a review.
| # | Control | What it stops | Evidence it produces |
|---|---|---|---|
| 1 | Inventory every server and transport per environment | Unknown and shadow servers | A registry row per server: owner, transport, version, scopes |
| 2 | Pin tool definitions by schema hash at approval time | Tool poisoning and rug pulls | A stored hash per tool plus a drift alarm on re-discovery |
| 3 | Scope credentials per tool, not per agent | Excessive agency and lateral reach | The identity each tool call actually used |
| 4 | Sandbox local and third-party servers | Local compromise and data exfiltration | Container and network policy applied to the server process |
| 5 | Require explicit consent for one-click server installs | Malicious startup commands | The command shown, and the approval recorded against a principal |
| 6 | Bind session identifiers to the verified principal | Session hijacking and event injection | The session-to-principal mapping, with the mismatch alarm |
| 7 | Enforce per-identity tool allowlists, default deny | Privilege escalation by scope creep | A deny row naming the tool, the identity and the reason |
| 8 | Re-validate on every version change | Silent capability growth | A re-approval record tied to the new definition hash |
Two properties distinguish this from a checklist that looks good on paper. Every row has an owner and a timestamp, so the list can be audited rather than admired. And steps 2, 6 and 7 fail closed: a definition that does not match its pin, a session presented by the wrong principal, or a tool outside the allowlist produce a recorded refusal rather than a warning somebody has to notice.
What those recorded refusals must carry to survive an evidence review — attribution, retention and an exportable format — is the discipline on AI compliance.
What to check: read the list as an attack plan. For each row, ask what an attacker who already controls one tool response would do, and whether the control produces a line you would see. Keep the pinned definitions where the review can read them, so the artifact the gateway enforces and the artifact a reviewer approves are the same one.
model context protocol security: transport, session, and authorization
The protocol-level surface has three layers, and each one has a specific default you should verify rather than assume. Transport. The Streamable HTTP transport lets a server assign a session identifier at initialization and hand it back in a response header; clients must then send it on every subsequent request, and a server that requires it should answer a request without one as a bad request rather than by creating an implicit session. Origin validation and binding to the local interface for local servers belong to the same section.
Authorization. MCP uses OAuth 2.1, and the security of the whole chain rests on audience validation and on refusing to relay credentials you did not issue. Two more requirements matter operationally: the authorization URL must use a safe scheme, and a client deployed to a server has to consider request forgery risk when fetching OAuth-related URLs, because during discovery it fetches URLs that a malicious server controls. The specification prescribes blocking private and reserved address ranges and routing discovery through an egress proxy that blocks internal destinations, and it flags the check-to-use gap: a name can resolve to a safe address during validation and to an internal one during the actual request, so the resolution should be pinned between the two.
Identity granularity. Scopes decide how much a stolen or misused token is worth. The recommendation is a progressive, least-privilege scope model: start from a minimal read-oriented set, elevate on a targeted challenge when a privileged operation is first attempted, and accept reduced-scope tokens rather than demanding the full set.
What to check: three concrete readings. The rejection reasons for Origin and Host validation, which should exist separately from authentication failures. The session table, which should show a principal column derived from a verified token for every live session. And the token's audience claim at the point of use — a service that trusts a token without checking the audience is the finding. Where the tokens are issued and renewed is covered in depth by MCP OAuth and auth; the transport and session rules above are what the protocol itself fixes.
mcp gateway security: one enforcement point, one log
If the permission boundary, the server allowlist and the credential scope are enforced in three different places, you have three places to be wrong and no single answer to "what can this agent do". Putting them in one enforcement point in front of the model is what makes the earlier sections operable — and it is a different thing from a network proxy, which sees connections rather than tool calls.
A gateway earns its place by owning four decisions at the moment of the call. Identity: which key, team and session the call belongs to, resolved before the tool runs. Scope: whether that identity is allowed to call this tool on this server at all, answered default-deny. Rate and budget: how much this key may spend per minute and per month, so a runaway loop is a refusal rather than an invoice. Record: one row per call — principal, server, tool, decision, latency and outcome — written where the agent cannot rewrite it.
How a call is priced against a team, a window and a monthly budget key is worked through with the shipped functions on AI financial analysis for agents.
Which layer owns those four decisions, where the per-key rate and the per-team budget are enforced once instead of per service, and how over-reach is noticed while it is happening are the control-plane questions on AI agent governance.
The property that makes this a security control rather than a convenience is fail-closed behaviour and separation of duties. The agent's own identity should not be able to write, delete or truncate the audit rows, because an agent that can edit its own record can hide its own actions — the same argument OWASP's MCP audit-and-telemetry entry makes when it asks for immutable trails. The deny path is equally important: a refusal that is not recorded cannot be alerted on.
What to check: exercise the deny path deliberately. Call a tool outside the identity's allowlist and confirm three things at once — the call did not reach the server, the refusal appears as a row with the identity, the tool and the reason, and the refusal is visible in the same view as the successful calls, so an operator does not have to reconcile two systems. Then confirm the agent's credentials cannot write that store. The record layer's shape is worked through on LLM observability; the enforcement layer we operate for this purpose is described on the MCP gateway page.
mcp prompts: the user-controlled primitive
Tools are chosen by the model; prompts are chosen by the person. That distinction is the whole design of the primitive, and it carries security consequences. In the specification, prompts are server-authored templates the client exposes for explicit selection, so the server declares a prompts capability, answers a list request, and returns rendered messages on a get request. The content is defined by the server; the choice of when to use it belongs to the user.
Three details in that design are security-relevant. Prompt arguments are a form a person fills in rather than a payload a model constructs — the specification gives them names, descriptions and a required flag instead of a full input schema. The advertised set may legitimately vary with the authorization presented on the request, so "this server shows me different prompts" is a feature rather than a compromise, but it also means the trusted set is per-caller. And a server that declares the list-changed capability can alter its prompts at runtime, which the specification says must not vary per-connection or as a side effect of another request but does allow to change over time.
The injection surface is therefore in the arguments and in the update path. A template that interpolates a raw argument directly into instruction text turns a user-supplied field into an instruction slot, so arguments should be validated where the prompt is rendered and passed as data with an explicit boundary. A prompt that changes after review is a change to your trusted instruction set: re-read it, re-approve it, and record who approved which version. The cheapest control is to treat prompts as reviewed artifacts rather than server configuration — they are instructions you are about to place in the user's turn, authored by someone outside your organisation.
What to check: log the render calls, not only the tool calls. A row per prompt retrieval — server, prompt name, argument keys, caller identity — answers two questions: which templates are actually in use, and whether a template changed between two runs of the same task. Then pass an argument containing instruction-like text and confirm it arrives as data in the rendered message rather than being executed. That test separates a parameterised template from an injection slot.
Where SmartGate fits
SmartGate is the enforcement point this page describes: the place where a tool call is attributed to an identity, checked against a scope, charged against a limit, and written into a record — the same object that the per-key rate limit and the per-team token budget read. That is the layer worth owning early, because the other controls above all consume its rows, and it is the one we operate.
The same server is what an individual connects to their own assistant for research, described from the user's side on SmartGate MCP for research and decisions.
The plan table sets the operational limits, not the features: monthly token caps of 2M, 20M, 100M and 200M+, MCP requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and 2, 10, 30 or 9999 keys per team. Reconcile those with your own retention and rate requirements before you commit a date to anyone. The pricing page is the authoritative table and takes precedence over anything quoted here.
Frequently Asked Questions
Limitations
This page is a control map, not a certification and not a benchmark. It names the attack classes the protocol and the major frameworks document, and the controls that address them; it does not claim that any product implements them correctly, and it does not rank vendors. Where a claim depends on a specific product's behaviour, verify it in your own environment rather than taking this page's summary as the documentation.
Two honest gaps. First, prompt injection has no complete mitigation today: the measured techniques reduce attack success substantially, and none of them make the class disappear, so the design goal is to bound the impact of a successful injection rather than to prevent every one. Second, the protocol's own security requirements are only as good as the implementation behind them, which is why every section above ends with something to read in a log or a test to run — a control you cannot observe is a control you do not have.
The plan figures quoted in this page are limited to the four facts published with the product, and they change with the plan; a retention or rate decision should read the current pricing table rather than this page. Nothing here is legal or regulatory advice. This page carries no code excerpt, and the reason is recorded in the Method note below.
Sources
- The Model Context Protocol specification — Security Best Practices: modelcontextprotocol.io/docs/tutorials/security/security_best_practices (confused deputy, token passthrough, session hijacking, local server risks, one-click configuration consent); Authorization: modelcontextprotocol.io/specification/2025-06-18/basic/authorization (audience validation, the ban on token passthrough, egress-proxy guidance, progressive scopes); Transports: modelcontextprotocol.io/specification/2025-06-18/basic/transports (session identifiers, Origin validation, local binding); Prompts: modelcontextprotocol.io/specification/2026-07-28/server/prompts (user-controlled model, capability declaration, argument shape, list-changed notification).
- OWASP GenAI — AI Agent Security Cheat Sheet (blast radius, least privilege, memory isolation) and MCP Security Cheat Sheet (schema pinning, sandboxing, audit expectations), plus A Practical Guide for Securely Using Third-Party MCP Servers (container isolation, just-in-time access, registry-based discovery).
- OWASP — Top 10 for LLM Applications (LLM01 prompt injection, LLM06 excessive agency and complete mediation) and the OWASP MCP Top 10 (tool poisoning, scope creep, shadow servers, audit and telemetry, context over-sharing).
- NIST — AI Risk Management Framework, for the govern / map / measure / manage structure a security review will map controls onto.
- NSA — Model Context Protocol: Security Design Considerations, for token lifecycle gaps and cross-server tool-name collisions.
- Microsoft — Prompt Shields (user-prompt versus document detection and the intervention points) and spotlighting; the research behind it is Defending Against Indirect Prompt Injection Attacks With Spotlighting.
- Anthropic — Mitigating the risk of prompt injections in browser use and the jailbreak and prompt-injection mitigation guidance.
- AWS Prescriptive Guidance — Best practices to avoid prompt injection attacks, for tag spoofing and separating instructions from retrieved content.
- Demand figures are our own measurements: DataForSEO Google Ads, United States, English, measured
2026-09-30, recorded in this project's
search_volume.jsonandresearch_brief.md. Product plan figures were re-verified against the pricing page on 2026-09-30.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of 8 sections for this page (0 abstention(s), 8 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the eight section keywords, because this lane's vocabulary — security, protocol, prompts — collides with generic helper, type and route names throughout a product codebase. A pinned generic name would have given the page the shape of a verified article with none of the substance, so every section above is written from the public sources listed above, which is the house rule for an unpinned section.
Each section follows the same three-part shape on purpose: what the attack surface is, which control addresses it, and what you can read or run to confirm the control is actually in place. Protocol claims come from the MCP specification's own authorization, transport, prompts and security sections; agent-risk claims come from the OWASP cheat sheets and Top 10 lists; the effectiveness figures for content marking come from the spotlighting research, and classifier behaviour from Microsoft's Prompt Shields documentation. The section keyword quoted above each heading comes from this project's own measured keyword set, not from a third-party tool. No internal hosts, credentials, batch identifiers or auction data are transcribed, so there is nothing here that has to be asserted verbatim.