SmartGate

MCP Proxy vs Router vs Gateway: Which Layer Decides What

An MCP proxy, an MCP router and an MCP gateway are three layers rather than three products. The proxy owns the connection: it terminates one MCP transport and opens another, and it decides nothing about whether a call should happen. The router owns the destination: it reads the request and decides which upstream server answers.

Short answer: An MCP proxy, an MCP router and an MCP gateway are three layers rather than three products. The proxy owns the connection: it terminates one MCP transport and opens another, and it decides nothing about whether a call should happen. The router owns the destination: it reads the request and decides which upstream server answers. The gateway owns the outcome: it decides whether the call is allowed, under whose identity, against which budget, and what the record says afterwards. With one upstream, one credential and no shared budget, the proxy is the whole stack.

Key takeaways

  • Three names, three decisions: connection, destination, outcome. A product that makes two of them is two components in one process, and should be documented that way before the first incident.
  • A pass-through that starts refusing calls because a team has spent its month is a gateway wearing a proxy's name. Renaming it is easy; the shared counter the rename requires is the real work.
  • The layer that can answer "who called?" is the layer that can be your audit; a byte pipe cannot, however long it keeps the bytes.
  • Add a second hop only when it makes a decision the first hop cannot.

Three words cover one neighbourhood, and demand does not sort itself by layer. In our own paid measurement (DataForSEO Google Ads, United States, 12-month window, 2026-09-30) "mcp client" carries 1,600 searches a month, "docker mcp gateway" and "streamable http" 880 each, "mcp proxy" 720 and "mcp router" 260. Those are not competing definitions of one thing; they are the vocabulary of different layers, and this page is about the line between them.

mcp proxy: the connection layer, and what it may decide

An MCP proxy terminates one MCP transport and opens another. That sentence is the whole remit, and most of the confusion around the name comes from adding to it. The job is reachability: the client and the server do not have to see each other directly, and they do not have to speak the same transport revision. Three situations justify one honestly: the server is a local process spoken to over standard input and output while the client is elsewhere; the client speaks the older two-endpoint HTTP transport while the server implements the current single-endpoint one; or the server sits inside a network the client cannot route to.

What the proxy decides is small by design. It decides which upstream a connection maps to, and in the ordinary case that mapping is one to one and lives in a configuration file. It decides how a client session identifier maps onto an upstream session, which is real work whenever the upstream keeps state and several clients share one endpoint. And it may normalise the envelope enough to make a transport-revision difference invisible to both ends — protocol knowledge, not policy.

What it cannot decide is whether the call should happen. The moment a pass-through refuses a tool call because a key exhausted its minute, or because this caller may not reach that server, it is no longer a proxy: it is a control layer that also forwards. That is not an argument about names — the rename is trivial, what it demands is not. A refusal by key or by team needs a counter that every instance shares, and a proxy deployed as a per-developer sidecar has nowhere honest to keep one.

One question decides how much a proxy can carry: does anything downstream know who called? If every caller arrives with the same static credential, then no component in the path can produce a per-caller limit or record, because the identity is not in the request. Authentication at the proxy is not authorisation by the proxy: a proxy can require a credential before it opens a connection, which is a door, not a judgement about who may walk through it. When that credential is issued per developer and the upstream treats it as its own, the upstream is the enforcement point and the proxy is still plumbing.

So when is one proxy enough? One team, one upstream, a credential per developer, no shared budget, no obligation to reconstruct who called what. Under those conditions a router is a lookup table with one row and a gateway enforces a policy that does not exist yet. The honest cost of that design is that the day you need the counter, you will add a layer rather than flip a setting.

mcp router: the first component that reads the request

A router answers exactly one question — which upstream should serve this call — and its output is a destination, not a verdict. To answer it, the router has to read the request envelope: the method, and for a tool call the tool name, matched against a registry of known servers. That is what makes the router the first component in the path that is not transport-blind. A proxy can forward a message it does not understand; a router cannot. A single binary that offers routing and policy together is therefore two components sharing a process and a configuration file, and it should be documented that way, because the first incident will otherwise be diagnosed in the wrong place.

The order of routing and authorisation matters more than most teams expect, and both orders ship. If the router runs first, an unauthorised caller learns which server names exist before anything refuses it, because the registry answers the question. If authorisation runs first, the router only sees calls that are already permitted and a refusal never touches a server. Only the second order lets you say that a refused call never reached an upstream — the sentence an incident review wants, and worth choosing deliberately rather than inheriting.

Name collisions are the second cost of aggregation. Two servers that both expose a search tool cannot both register the same name, so something must namespace them, and the cheapest mechanism — a prefix on every tool name — becomes public API the moment a client's tool list depends on it. Renaming a server behind that prefix is then a breaking change for every client whose configuration enumerates tools. A router that merges many servers behind one endpoint is not a fancier proxy; it is an interface you have to keep stable.

Failover is the third cost. A router that retries on a second upstream must know whether repeating the call is safe: for a read that is a design question, for a write a correctness question, and the answer lives in the tool's own contract rather than in the router's registry.

What the router does not decide is who may call and how much they may spend. It has no counter, and in a typical deployment no per-caller identity either: it sees a tool name, not a person. With exactly one upstream, the router collapses into a configuration file with one row plus a lookup that costs a hop, so the honest moment to add it is the arrival of the second upstream, not before. A router that resolves server names is still not a router of models, though, and the two jobs sit in different layers: the OpenRouter alternative comparison separates them.

docker mcp gateway: what the container shape adds, and what it cannot hold

This phrase names a shape more than a layer, and the search results agree: the leading answers are the project's own repository and documentation rather than explanations of a category. The shape is a container-local gateway — a command-line plugin that starts MCP servers as containers on a developer's machine and exposes them to a client through one endpoint.

What that shape is good at is local trust. Each server runs in its own container instead of as a child process of an editor, so one that misbehaves is bounded by the runtime rather than by hope. Secrets can be injected at launch time, so the client's configuration file does not have to hold an upstream token. And a curated catalogue turns a per-laptop choice into a review that happens once.

What the shape cannot be is the team's single enforcement point, and the reason is arithmetic rather than architecture. Every developer runs a copy, so a per-team token budget has as many counters as there are workstations, and a per-minute request limit resets whenever a machine sleeps. The same is true of the record: if each copy writes its own log, the audit exists in as many places as there are installations, and reconstructing it is a collection job rather than a query. A local instance can ship its rows to a central store — at which point you are operating a collector, and the decision you wanted centralised has simply moved somewhere you now have to run.

The pattern that works is two tiers with a deliberate split. The local tier decides which servers this developer's client sees, and holds the launch-time secrets, because both of those are per-developer facts. The central tier decides whether a call happens and keeps the one record. The rule for the split is worth stating plainly: a decision that needs a shared counter — a monthly token budget, a per-minute request limit — must be made where the counter lives, and a decision that needs the developer's local context is better made locally.

The container shape is also a supply-chain answer, not only a dispatch one: running a stranger's MCP server means running their code with your credentials nearby, and isolation is a decision about trust rather than routing. A container-local gateway improves that and leaves the shared counter open; a proxy with the same reachability responsibilities leaves both open.

mcp client: the layer that owns the model's tool list

Everything above sits on the wire between a client and a server. The client is the end of that wire, and three decisions belong to it that no intermediary can make on its behalf.

The first is which tools the model is offered. If a gateway aggregates many servers behind one endpoint, the client sees their union, and every tool definition is context the model pays for on every turn whether or not it is used. A control layer can help by exposing a different tool set per credential — a real lever, and the only one it has. Choosing how much of that list a given task needs is a client-side judgement about the conversation in front of it, and no proxy can make it.

The second decision is approval. A client is the only layer in the path that can ask a human a question before a call leaves, so a requirement such as "confirm before any tool that writes" has to be implemented there. A gateway cannot substitute for it: it can refuse, and it can require a different credential, but it cannot wait for a person. Teams that assume their policy layer provides human approval discover the gap the first time an unattended agent runs a write tool at three in the morning.

The third is the transport the client speaks: a local process over standard input and output, or a remote endpoint over HTTP. That single choice is what creates the proxies the first section describes. A client that can only launch local processes cannot reach a shared server, so something must be started on the developer's machine or in front of the network, and that something is the hop the rest of this page is about.

The client is also where the limits of what it cannot enforce show up. Each host keeps its own configuration file, so "one endpoint, edited in one place" is a different claim from "one line in each of five files". And the client cannot enforce a shared budget: two laptops with the same limit are two limits, so anything that must be one number for a team has to be true somewhere the client is not. The smallest complete example of those client-side decisions is connecting OpenClaw over MCP: one server entry, two headers, and the tool list the model is offered written down in one place.

streamable http: the transport is the proxy's business, not the policy's

The current MCP HTTP transport is one endpoint. The client posts requests to a single URL, and the server answers either with one response or with a stream of server-to-client messages, using the stream only when it has something to push. The revision before it used two endpoints — a long-lived stream plus a separate post path — and that difference is the most common honest reason a proxy exists in a 2026 deployment: one end of the connection was written against the older shape.

Which layer the transport belongs to is the point. Transport is the connection layer's business; policy is not. That matters because a control layer must terminate the same transport anyway — it has to read the request envelope to decide anything — so putting a proxy in front of it terminates the transport twice. The costs of that second termination are concrete and unglamorous: two session maps that must agree, two timeout settings that must agree, two places where a resumable stream can die, and one more hop where an idle connection can be reclaimed by something in the middle. None of those is fatal. Each is a component that must be configured and then owned.

The outer hop earns its place in a small number of situations, and they are all about reachability rather than policy. The inner endpoint is not routable from the client's network. The client cannot be configured with the headers or the path the inner endpoint needs. Or TLS has to terminate at a specific network edge and the same place must be the only ingress. If the same credential crosses both hops and the inner endpoint is reachable from the client, one of the two hops is dead weight — it will hold a session, consume a connection and make latency slightly worse without making a single decision the other hop could not make.

Session handling is where a proxy can still be load-bearing rather than decorative: if the upstream keeps per-session state in a process, a load balancer needs stickiness, and holding that mapping in one place is a genuine connection-layer job. The rate limits people worry about when they add a hop are unaffected by it — requests per minute per key are 120, 300, 600 and 1200 across the plan table, and a proxy does not change the count, because a proxy does not read a key. Only the layer that holds the credential and the counter turns that number into an enforcement.

mcp server security: three questions the connection layer cannot answer

A security review of an MCP deployment asks three questions, and they map onto the three layers cleanly. Does the server know who is calling? Is every tool call recorded with the caller's identity? Can a call be refused before it reaches the server? A transport-only proxy answers none of them, and that is a property of the layer rather than a defect of an implementation.

Four risks follow from that mapping, and each one lands on a different layer.

Server substitution. A router resolves a tool name to a destination, so whatever holds the registry decides which server the caller actually talks to. If that mapping is a file nobody reviews, a registry edit is an attack surface: neither the client nor the proxy can tell that search is now answered by a different process than it was yesterday. The fix is a versioned, reviewed registry with one namespace per server, so a substitution changes a name rather than silently changing a meaning.

The confused deputy. Forwarding the caller's own credential upstream means many callers arrive as one identity, so revocation stops being a small operation and the upstream's log cannot attribute what it did to a person. Holding the upstream credential in the control layer and passing per-caller identity in the record, rather than on the wire, keeps both possible — a design decision made at the gateway, not at the pipe.

Shared static credentials. If every developer holds the same token, no layer can name a caller and rotating it breaks everyone at once. This is the first section's question seen from the security side: identity in the path is the precondition for attribution anywhere.

Untrusted tool results. Tool output is input from somewhere you do not control, and a proxy cannot judge meaning, so the defences live where the content is consumed and where the consequences are bounded: what the model may do next, and which tools the credential in play can reach. A layer that filters tool sets per key does more for this risk than a layer that forwards faster.

One transport-level hygiene rule belongs here because it is nobody else's job. Retention of whatever the path records is a policy decision with a plan attached — audit-log retention runs 7, 30, 90 or 180 days depending on the tier, and the shortest is a deliberate limit rather than an accident. A pass-through that keeps request bodies with no owner and no retention decision is a liability that grows quietly. And the layer's honest limit stands: a proxy that logs bytes without a caller identity cannot produce the record "who called which tool", however long it keeps them.

Four questions before you draw the diagram

Written as questions, because the answers decide how many components you actually need.

  1. How many upstream servers are there? One means no router: the mapping is a configuration file, and adding a resolver buys a lookup and a hop. Two or more means something must resolve names, and that something becomes an interface the day a client's tool list depends on it.
  2. Does anything downstream need to know who called? If an audit, a per-developer limit or a revocation that names one person is a requirement, then a component must hold per-caller identity, and a transport-only proxy cannot be that component.
  3. Where does the counter live? A monthly token budget and a per-minute request limit are shared state. The layer that enforces them has to be the layer that owns them, which is why a per-developer sidecar can hold a catalogue and a secret but never a team budget.
  4. Does a human need to approve anything? Only a client can ask, and only the client can wait for the answer. A gateway can refuse a call; it cannot hold a conversation.
The decision Which layer owns it What that layer must be able to see
Can this client reach that server at all? connection layer the transport and the network, not the request
Which upstream serves this tool name? router the request envelope and the server registry
May this credential spend another token? control layer a shared counter and a per-caller identity
May this call run without a person? client the conversation, and a human

The collapse rule is short. One upstream and no shared budget: run the connection layer and stop. Several upstreams behind one credential that a team shares: you have already built the control layer, whether or not you named it one, and the counter is the part that will be missing. The mistake to avoid is adding the middle layer to postpone the last question — a router with no counter and no identity cannot answer the two questions that mattered.

Where SmartGate fits

This page is about the line between three names; where our product sits on that line is the control layer. Every tool call through the gateway is written as an audit row — caller, route, transport, tool, token count, latency, outcome — and that row is what the per-key rate limit and the per-team token budget read, so the record and the enforcement point are one object rather than two systems that have to agree. A proxy in front of it would forward bytes with no counter; a router behind it would choose among destinations the caller was already allowed to reach.

The plan table is a set of operational limits rather than a feature list: monthly token caps of 2M, 20M, 100M and 200M+; requests per minute per key of 120, 300, 600 and 1200; audit-log retention of 7, 30, 90 or 180 days; and 2, 10, 30 or effectively unlimited keys per team. If your retention requirement comes from an audit or a compliance deadline, compare those numbers with the requirement before you promise a date to anyone — the pricing page is the authoritative table, and you can start free to read your own limits on real traffic.

For the rest of the decision, the neighbourhood has its own reading order. If you are choosing between layers at the design stage, the architecture questions live on enterprise AI gateway architecture; if the constraint is that nothing may leave your network, begin with deploying a gateway in a private cloud; the comparison with the proxy most teams already own is drawn on AI gateway vs API gateway; and what a gateway is, as opposed to which layer it is, is answered on MCP gateway. If the pass-through already in production is a self-hosted proxy you are outgrowing, the criteria for leaving it - protocol, credentials, record, limits, carry-over - are on the LiteLLM alternative page.

Frequently Asked Questions

Limitations

This page is a boundary map, not a product review. It does not rank proxies, routers or gateways, because the right split depends on where your counter, your credentials and your record already live; it names the line between the three names and the question that moves the line, and it makes no claim that any named implementation belongs cleanly to one layer.

Naming in this ecosystem is loose: several projects take the name of a layer they only partly implement, and the container tooling cited above is one example — the shape is container-local, while the layer it participates in is the connection layer.

Nothing here is a security guarantee. A layer diagram says where a control can be enforced, not that it is configured, tested or reviewed: each risk named above needs its own test and its own owner, and the right boxes are not evidence that either exists.

The plan figures are operational limits read from the pricing page, not a feature comparison, and they change: they are no substitute for the current table before committing a retention requirement.

This page carries no code excerpt, and the reason is recorded in the Method note below.

Sources

  • The Model Context Protocol specification, transports — modelcontextprotocol.io/specification, for the current single-endpoint streamable HTTP transport and the earlier two-endpoint form this page compares it with.
  • The Model Context Protocol client concepts — modelcontextprotocol.io, for what a client owns: the tools it offers, the roots and sampling features, and where the conversation state lives.
  • Docker's MCP Gateway documentation — docs.docker.com, and the project repository at github.com/docker/mcp-gateway, as the container-local shape described above; Docker's own announcement at docker.com states the same isolation motivation.
  • Demand figures quoted in this page are our own measurements: DataForSEO Google Ads, United States, 12-month window, measured 2026-09-30, recorded in this project's search_volume.json and research_brief.md.
  • The plan table and the retention figures: taken from this project's brief, which records them as re-verified against the live /pricing page on 2026-09-30. The pricing page is authoritative and is the place to check before a commitment.

Method note

This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of 6 sections for this page: rule A found no unique symbol for any of the six section keywords. Five of them returned several equally plausible candidates with no rule-based match; the sixth matched one candidate whose name is a bare generic property, for which the slot proof returned no asset, so it was abstained rather than pinned. It recorded 1 abstention(s) and 5 no-slice verdict(s), and a pinned generic name would have given the page the shape of a verified article with none of the substance; the house rule for an unpinned section is to write it from sources, which is what every section above does.

Product limits came from this project's brief, which records them as re-verified against the live pricing page on 2026-09-30, and the heading vocabulary comes from this project's own paid measurement run rather than a third-party tool. No code, batch fingerprints, auction data or internal hosts appear here, so nothing on the page has to be asserted verbatim.