Agent Sandboxing: Execution Isolation for Tool-Using Agents
Agent sandboxing is the execution boundary between a model's decision and the machine that carries it out: a container, a microVM or a user-space kernel that runs tool-authored code with its own filesystem, its own network and the least privilege the task can survive.
Short answer: Agent sandboxing is the execution boundary between a model's decision and the machine that carries it out: a container, a microVM or a user-space kernel that runs tool-authored code with its own filesystem, its own network and the least privilege the task can survive. It matters because the one tool that executes code is the capability that turns a wrong token into a compromised host, and no instruction in a prompt can replace a boundary the code cannot negotiate with.
Key takeaways
- Execution is the highest-risk tool class. Reading is bounded by what you hand over; running code is bounded by what the host allows, so the strength of the boundary — not the wording of the system prompt — is the real control.
- Isolation is a spectrum, not a switch. A shared-kernel container, a user-space kernel and a hardware-virtualised microVM sit at different points between "convenient" and "hard to break", and the right point depends on the trust you place in the code.
- The boundary belongs around the call, not the process. Filesystem scope, network egress, resource ceilings and credential injection are decided per tool invocation; a sandbox that is opened once and reused leaks state between tasks.
- Escapes are ordinary systems bugs. Shared kernels, mounted control sockets, reachable cloud
metadata endpoints and
procfilesystems are the recurring shapes; sandboxing reduces the blast radius rather than promising a wall. - Isolation has a bill. Cold starts of tens to hundreds of milliseconds, missing filesystem features and blocked network calls are the price, and it is paid on every task, not once.
- Decide the boundary before you pick the vendor. Write down what the code may touch, then measure candidate runtimes against that list.
agent sandboxing: the boundary that runs the code
An agent is a loop that decides, then acts, and most of what it decides is reversible: a search that returns the wrong page, a summary that misses a clause. One tool class is different. When the agent can execute code — a shell, a Python interpreter, a build step, a database client that runs a statement — a wrong decision has a side effect on a real machine, and the size of that side effect is set by the privileges of the process that runs it. That is the problem agent sandboxing exists to answer.
The reflex is to describe the boundary in the prompt: "only read files under /data", "never touch the network", "do not run destructive commands". These are instructions, and instructions are input to a probabilistic system. They raise the cost of a mistake; they do not stop one. A boundary the code cannot argue with is enforced somewhere the model has no reach — a namespace, a hypervisor, a syscall filter, a kernel that never sees the host. The distinction is between telling the agent what to do and constraining what it is physically able to do, and only the second is a control. An unfenced code tool is the payload end of the prompt injection attack surface, which is the centre page of this neighbourhood.
Three properties make execution different from every other capability. It is composable — the agent can write code that calls code, so a single injection can become a script. It is open-ended — a shell is not one action but a programming language over the host's entire surface. And it is stateful — a file written, a process started or a socket opened this turn changes the world the next turn runs in. A read tool that returns junk produces a bad answer; an execute tool that is loosely bound produces an incident.
agent sandbox: four isolation levels and what each one stops
"Is it sandboxed?" is the wrong question, because the word covers mechanisms with very different strength. The useful question is what runs the untrusted code, and what can it still see. Four levels cover almost everything in production, and they differ in a property that is easy to state and hard to fake: whether the untrusted code shares a kernel with the host.
| Level | What isolates the code | Stops | Cold start | Compatibility |
|---|---|---|---|---|
| Process / language runtime | a restricted interpreter, no ambient credentials | accidental damage, not a determined attacker | microseconds | Low — the safe subset is smaller than the language |
| Shared-kernel container | namespaces and cgroups on the host kernel | most cross-task reads, resource exhaustion | tens of ms | High — a near-normal Linux userland |
| User-space kernel | a re-implemented kernel between the app and the host (gVisor class) | most host-kernel syscall surface | ~100 ms | Medium — some syscalls and io_uring gaps |
| Hardware microVM | a virtual machine boundary (Firecracker class) | host-kernel bugs; the guest has its own kernel | ~125 ms to boot | Medium-high — a full kernel, minus nesting and passthrough |
The reading of that table is the whole discipline. A shared-kernel container is the default
because it is fast and familiar, and it is a resource and access boundary, not a security wall: a
host-kernel vulnerability is reachable from inside it, and a mounted control socket (docker.sock
and friends) hands the container the host back. A user-space kernel interposes on the syscall
surface, so the untrusted code talks to a re-implementation that returns plausible answers, and the
interposition itself is the attack surface. A microVM gets its own kernel, which is why it
survives the class of bug that defeats a shared kernel, at the cost of boot time and of some
hardware features that need passthrough. The engineering job is to pick the cheapest level whose
weakness you can actually tolerate, and to write that choice down.
ai agent sandbox: the boundary around one tool call
The second decision is scope: one sandbox per agent process, or one per tool call. Per-process scope is simpler and cheaper, and it is wrong for an agent. A long-lived sandbox accumulates state — files from a previous task, environment variables, a warm connection, a cached credential — and the next task inherits all of it. The boundary that matches the agent's actual unit of work is the call.
A call-scoped sandbox is a small envelope built for one invocation and destroyed after it. Six fields define the envelope, and each is a decision rather than a default:
- Filesystem — a writable scratch directory plus an explicit allowlist of mounts. The code should not see the host's home directory, the agent's own configuration, or any credential store. If it needs one input file, copy that file in; do not mount the tree it lives in.
- Network — default deny, then allow the endpoints the task genuinely needs. A code tool that can reach the whole internet can exfiltrate everything it can read, and can reach the cloud metadata service, which is a credential source on most hosts.
- Privilege — a non-root user, no privileged flags, a read-only root filesystem where the runtime allows it, and no inherited capabilities.
- Credentials — injected per call, scoped to the task, and short-lived. Ambient secrets in the environment are readable by any process in the sandbox, so a tool that only needed a read token should never be handed the write token "just in case".
- Resources — CPU, memory, wall-clock and file-size ceilings. An escape is one failure mode; a fork bomb or a disk-filling loop is another, and it needs no bug — only a model that wrote an unbounded loop, which also makes the ceiling the stop condition for a runaway turn.
- Cleanup — the envelope is destroyed, not paused. A sandbox that survives its call is a cache of the last task's state, and the next task reads it.
Where a tool can be split, split it. A tool that reads cannot exfiltrate what it cannot see; a tool that writes to a scratch space cannot damage a production database. The breadth of what a task needs is almost always narrower than the breadth of what it asks for, and the job of the boundary is to hold it to the narrower one.
sandboxed code execution: the thin wrapper at the boundary
Every sandboxed executor, whatever it is built on, arrives at the same shape: a method whose whole
body hands a call to the primitive that actually runs, and returns the result. The excerpt below is
that shape, not a sandbox in itself — a Redis pipeline's exec() method, which does nothing but
forward to the underlying client. It is worth reading because it is the pattern a boundary method
should have: one job, no policy of its own, and no place to hide a second path to the host.
# lib/redis/upstash-adapter.ts — source lines 49–51 (the exec boundary a tool call is handed to)
async exec() {
return p.exec();
}
The lesson is in what is absent. There is no argument parsing, no fallback, no "if the sandbox
fails, run it anyway" branch. An execution boundary is the worst possible place for a convenience
path: a single catch that falls back to the host, a
single debug flag that disables the envelope, or a single "internal" caller that skips the check
converts the whole control into a suggestion. The method forwards; the runtime, configured outside
the method and outside the model's reach, does the isolating. Keeping that split clean is the entire
security argument, and it is the reason the boundary is owned by the platform team rather than by
the code the agent writes.
Two consequences follow for the calling side. First, the wrapper must be the only route — if the agent can reach the underlying runtime directly, the boundary is bypassed. Second, the wrapper's return value is data, and data from an untrusted execution is untrusted data: a program can print a prompt injection, a credential, or a 200 MB blob, so the result is size-capped and treated as unverified input before it re-enters the model's context.
sandbox escape: the shapes that break the boundary
A sandbox escape does not usually begin with a clever model. It begins with an ordinary systems weakness that was reachable from inside the envelope, and the recurring shapes are worth naming because each names a check missing from the boundary.
A shared kernel. A container is not a kernel boundary, so a host-kernel privilege-escalation bug is reachable from inside it, and history keeps producing that class. A user-space kernel or a microVM moves the same bug out of reach.
A mounted host surface. The control socket (docker.sock), a host device node, a
/proc interface or a shared volume mounted "for convenience" all hand back a path the sandbox was
meant to remove. Any mount of a host-managed path is a hole; the whitelist should be small enough
that each entry can be justified.
A reachable network. A sandbox with free egress can reach the cloud instance-metadata service, an internal API, or the agent's own control plane, and can exfiltrate whatever it read on the way in. Network policy is part of the boundary, not a separate concern.
A missing ceiling. A recursive fork, an unbounded file write or a decompression bomb does not escape anything — it just exhausts the host from inside. Resource limits are a limb of the same control.
A side channel. Timing, cache and speculative-execution channels survive many logical boundaries, a reason not to co-locate one tenant's secret with another tenant's code.
The honest conclusion from the list is that sandboxing is blast-radius reduction, not a promise. Each level removes a family of reachable paths and leaves the rest; the goal is to make the remaining paths require a bug the attacker has to find, rather than a flag the code can set. That is also why the boundary is layered with the other controls in this cluster — detection, policy and identity bound what a successful escape can still do — rather than relied on alone.
llm sandbox: running model-authored code without trust
An LLM sandbox runs code the model wrote during a turn, and the model is a first-class untrusted author. It has read material you did not curate, it optimises for the task you gave it rather than for your infrastructure, and it can write a program that looks reasonable and does one extra thing. The difference from ordinary untrusted code is volume and intent: it arrives continuously and is generated to satisfy an objective rather than deployed by a reviewer. Language-level isolation is the mechanism most closely matched to that workload.
A language-level sandbox runs the code inside a restricted interpreter rather than on a Linux host. A WebAssembly runtime, a capability-restricted runtime such as a Deno-style one, or an in-browser Python runtime all execute a smaller, well-defined instruction set and simply do not expose the syscalls, the filesystem or the sockets a script would use to escape. The code is given a handful of explicit host functions — read this, fetch that — and nothing else exists. That is why these runtimes start in microseconds and why their weakness is the opposite of a container's: instead of being strong-but-leaky, they are small-and-closed, and the attack surface is the host-function interface you wrote. Everything you did not expose is unreachable by construction.
Two rules keep a language-level sandbox honest. First, the host functions are the boundary, so each
one is reviewed like an API grant — a fetch that resolves any URL is a network hole, and a
filesystem read that takes any path is a filesystem hole. Second, the output is untrusted text:
whatever the program prints returns to the context, so it is length-capped and clearly labelled as
data, because a program the agent wrote is an excellent delivery vehicle for a second-stage
instruction. When the code needs a full userland — native binaries, system packages, a real shell —
the language-level approach runs out and the host-level levels above take over. The choice between
them is the choice between "small and closed" and "large and interposed", and it is decided by the
dependencies the task actually needs, not by preference.
mcp sandbox: isolating a server you did not write
An MCP server is a process that exposes tools to the agent, and it is frequently third-party code run on your infrastructure with your credentials in its environment. That combination is its own trust boundary. The agent decides to call a tool; the server decides what the call does, and a server you did not write is untrusted code holding a tool description the model will believe.
The isolation discipline for a server is the call-envelope discipline from above, applied to the server rather than to a snippet: run it as its own process or container with no ambient access to the agent's filesystem, inject only the credentials its declared tools need, and set an egress policy that matches those tools. A filesystem server gets a scoped directory; a search server gets the search endpoint and no write credentials; a server whose tools are read-only should not be running as a user that can write. The common failure is the reverse: one long-lived server process, started once with a broad token in its environment, shared by every task, so any task can reach every capability the token holds. Splitting the server per trust level is the cheap fix, and it is the same split the MCP gateway layer performs at the protocol boundary.
Two MCP-specific details are easy to miss. First, a tool description is model input, so a malicious server can inject instructions through its own tool metadata — the boundary must treat server-provided text as untrusted even when the server itself is sandboxed — a propagation path this page leaves to the indirect-injection work in the same cluster. Second, the server lifecycle is part of the boundary: a server that outlives the task that needed it carries context forward. Call-scoped servers cost more to start and leak less, and for third-party code that trade is usually worth making.
How SmartGate's gateway sits above the sandbox
SmartGate is not a sandbox, and it does not claim to be one. It is the layer above the boundary that decides which tool calls happen at all: per-key rate limits, per-team token budgets, and an audit row per call. Read against the isolation levels above, the two layers compose in a specific order — the sandbox bounds what one execution can do, and the gateway bounds how many executions a caller can start and records what each one did. Neither substitutes for the other: a sandbox without a call record cannot tell you which task caused an incident, and a gateway without a sandbox leaves the executed code with the host's privileges.
The plan table sets operational limits rather than boundary strength: monthly token caps of 2M, 20M, 100M and 200M+, requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and up to 2, 10, 30 or unlimited keys per team. Those numbers belong to the control layer — they cap the traffic the sandbox is asked to run and keep the record long enough to review it — and they move with the plan, so the pricing page is the authoritative table.
What our fetcher refuses, read from our own egress checks
"Sandboxing" for an agent is often used to mean two different things: constraining what code can do,
and constraining where a tool is allowed to reach. A gateway can only do the second, and it should say
so. Here is ours, read on 2026-10-07 from backend/smartgate/modules/fetch/algorithm.py.
- Only two schemes are allowed. Everything else — local files, FTP, data URLs, and a handful of
historical protocols — is refused outright. Allowing a fetcher to accept
file:is how a "fetch a URL" tool becomes a file reader. - The hostname is resolved and the resolved address is checked, against a list of non-routable ranges: the unspecified and loopback blocks, private ranges, and the link-local block that cloud metadata services live on, for both IPv4 and IPv6. Checking the resolved address rather than the typed hostname is the difference between blocking an obvious name and blocking the mechanism.
- Responses are capped at two megabytes, and results are cached briefly (two minutes, 128 entries, both adjustable). The cap protects the process; the cache means a chain that fetches the same URL twice does not pay twice.
- The distinction to keep straight: this is network egress policy, not a sandbox. It stops a tool from reaching your internal network; it does not limit what the tool's code does, what it keeps in memory, or which credentials it can read. A vendor page that blurs those two claims is asking you to accept the weaker guarantee under the stronger word.
- What it does buy you: an agent that can be pointed at user-supplied URLs without becoming a server-side request forgery primitive — which is the concrete, testable version of "sandboxed".
How to get started
- List what the code may touch, before choosing a runtime. Write the filesystem paths, the endpoints, the credentials and the ceilings a typical tool call needs. That list is what a candidate runtime is measured against.
- Pick the cheapest level that satisfies the list. If the task needs a shell and native packages, that is a microVM or a container; if it is a data transform over an API, a language-level runtime is smaller and faster.
- Scope the envelope to the call, not the process. Build the sandbox per invocation, inject exactly the credentials that call needs, and destroy it. Verify by starting two tasks and checking that the second cannot see the first's files.
- Wire the call record in the same release. The audit row is what turns "the sandbox ran something" into "this task ran this tool with these arguments". Start on the free platform tier to see the tool traffic and the per-key limits on real data with start free, then confirm the current limits against your own workload.
Frequently Asked Questions
Is a container the same as a sandbox?
No, and the difference is the kernel. A container is a namespaces-and-cgroups boundary on the host kernel, which is fast and familiar but shares the kernel with the host; a host-kernel privilege-escalation bug is reachable from inside it. A sandbox in the stronger sense — a user-space kernel or a microVM — gives the code its own kernel or interposes on the syscall surface. A container is a resource and access boundary; treat it as one.
What is the smallest thing that actually stops a determined escape?
A boundary that does not share a kernel with the host, plus a network policy that denies egress by default. A microVM gives the untrusted code its own kernel; the egress policy stops it reaching a metadata service or an internal API even if it runs code successfully. Either alone is incomplete — a strong boundary with open network can still exfiltrate, and a closed network on a shared kernel can still be broken.
Does sandboxing slow the agent down?
Yes, and the cost is the cold start. A shared-kernel container starts in tens of milliseconds, a user-space kernel around a hundred, and a hardware microVM around 125 ms to boot before the code runs. Paid once per call, that is the number that decides whether a per-call envelope is affordable for a chat turn. Warm pools and snapshot-resume cut the latency, at the cost of holding idle capacity.
Should the sandbox be reused across turns of one task?
Within a single task, reuse avoids paying the cold start on every turn and lets the code keep its scratch files. Across tasks, never: a reused sandbox carries files, environment variables and warm connections into the next task, which is the leak the boundary exists to prevent. The rule that satisfies both is scope the sandbox to the task, destroy it at the task boundary, and make the cold start a measured cost of the design.
Where does the sandbox end and the guardrail begin?
They are different controls answering different questions. The sandbox bounds what code can do to the host — a property of the runtime, enforced outside the model; a guardrail inspects inputs and outputs for bad content — a probabilistic filter applied to data. A prompt-injection attempt can pass every guardrail and still be harmless because the sandbox removed the capability it wanted, and a perfectly isolated sandbox can still leak data through what it is allowed to return. The patterns that catch the second case belong to the LLM jailbreak and LLM security pages in this cluster.
Limitations
This page describes a boundary, and a boundary is only as strong as its configuration. The isolation levels above are classes, not products: a well-configured container can out-perform a badly configured microVM, and no level removes the need to review the mounts, the egress policy and the credentials it is given. Nothing here is a ranking of vendors, and nothing here is a guarantee that a given runtime has no escape — a runtime is software, and software has bugs.
Isolation also does not make untrusted content safe. A sandbox can stop executed code from reading a credential or reaching the network; it cannot stop the model from being persuaded by a poisoned document it read, because that failure happens before any code runs. It needs the input and detection controls of the security-framework work and the identity scoping of the agent-identity work alongside the sandbox, not instead of it.
Finally, the costs quoted are the costs of the mechanisms, and they move with implementation. Cold starts improve with snapshotting; compatibility gaps close as runtimes add syscalls; and a boundary that is affordable today can be re-priced by the next hardware generation. Treat the numbers as shapes to plan against and re-measure them on your own workload.
Related reading in this cluster
This page owns execution isolation. The pages below own the adjacent questions, and the boundary above is built to compose with them: the centre page frames the attack surface, and the others own detection, identity, the propagation path, the model-level corollary and the platform.
- The prompt injection attack surface — the centre page: why a tool chain turns injected text into a real action.
- An AI security framework's control mapping — where governance requirements land as testable controls.
- Scoping agent identity and credentials — who the agent is, and what a call is allowed to prove.
- How indirect prompt injection propagates — the path from untrusted data to instruction.
- LLM jailbreak patterns — the model-level corollary.
- Operating LLM security — keys, logs, retention and deployment.
Sources
-
gVisor documentation — gvisor.dev/docs, for the user-space kernel model, checked 2026-10-04.
-
Firecracker — firecracker-microvm.github.io, for the microVM model and boot latency, checked 2026-10-04.
-
OWASP Top 10 for LLM Applications — genai.owasp.org, for the excessive-agency and supply-chain entries, checked 2026-10-04.
-
Kubernetes SIG
agent-sandbox— agent-sandbox.sigs.k8s.io, for sandbox lifecycle and network-policy patterns, checked 2026-10-04. -
WebAssembly and the WASI capability model — wasi.dev, for language-level isolation, checked 2026-10-04.
-
Demand figures for this page's section keywords are our own measurement: DataForSEO Google Ads, United States, English, 12-month window, recorded in this project's
search_volume.jsonandresearch_brief.md. -
Product behaviour and the plan table: read read-only from the product source at the revision recorded in this project's
pipeline_results.json, with the plan figures re-verified against the live/pricingpage on 2026-10-04. -
The section "What our fetcher refuses, read from our own egress checks" is our own implementation, read on 2026-10-07 from
backend/smartgate/modules/fetch/algorithm.py(origin/main). It states only what those files state.
Method note
This page carries one code excerpt out of seven sections, and that split is the honest result of the
slice matcher, not an editorial choice. pipeline.alethix run pinned 1 of
7 sections for this page, with 0 abstention(s) and 6 no-slice
verdict(s). Rule A found a unique symbol for the sandboxed code execution section at level 2 — the
symbol exec is a token substring of the keyword's execution, and it is the only local match — and
a server-side proof call confirmed it, so that section carries a fenced block cut verbatim from the
slice body. The other six section keywords returned either no local match or a set of equally
candidates that could not be uniquely resolved (generic helpers such as Memory.delete_all that say
nothing about sandboxing), so those sections were written from external sources, cited above with a
checked date. A generic symbol pinned there would have given the page the look of verified code with
none of the substance, and the house rule is that an unpinned section is sourced, never invented.
The pinned excerpt is a thin forwarding wrapper — a Redis pipeline's exec() — and the section says
so plainly rather than presenting it as isolation code; it is quoted because it is the shape every
execution boundary shares, not because it proves a sandbox. Product claims were read read-only from
the product source at the revision the slice run recorded, and the plan figures were re-verified
against the live pricing page on 2026-10-04. Demand figures come from this project's own paid
measurement run. No code is transcribed by hand, no batch fingerprints, auction data or internal
hosts appear in the text, and the fenced block above was re-asserted byte-for-byte against the slice
body before publication.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|---|---|---|---|---|---|
| 1 | sandboxed code execution | exec |
lib/redis/upstash-adapter.ts |
49–51 | rule A L2 → slot-proof | 0d15c81fb7ea |