SmartGateSmartGate

What Is an MCP Gateway? Roles, Boundaries, and When You Need One

An MCP gateway is a single authenticated entry point that sits in front of one or more Model Context Protocol servers, so every agent host points at one address instead of a list of them. It terminates the session and answers the caller-side questions — who is asking, which server should receive the request, what policy applies, and what gets recorded — before any tool server sees a call.

Short answer: An MCP gateway is a single authenticated entry point that sits in front of one or more Model Context Protocol servers, so every agent host points at one address instead of a list of them. It terminates the session and answers the caller-side questions — who is asking, which server should receive the request, what policy applies, and what gets recorded — before any tool server sees a call. It is not the tool server, and it is not a byte-forwarding proxy: the server owns the tools, and only a component that reads the MCP call can enforce a limit per key or per team. The rest of this page is the framework behind that sentence: the four jobs, the boundary with the server, the boundary with the model-facing "AI gateway", and when a gateway is worth adding at all.

Key takeaways

  • The gateway owns the caller; the server owns the tools. A raw MCP server answers tools/list and tools/call; it carries no position on who is allowed to call, how often, or at what cost.
  • Four jobs define the layer — authentication, routing, policy and audit — and a component that skips the policy and audit jobs is a proxy wearing the word "gateway".
  • The boundary is what it parses. A proxy sees bytes and a destination; only something that reads the MCP message can meter it per key, per team or per month.
  • Model-facing gateways are a different layer. An "AI gateway" for model providers and an "MCP gateway" for tool servers route different traffic; the shared vocabulary is the most common cause of a wrong purchase.
  • Decide from what you cannot see. List the callers, servers and records your current setup cannot account for; whatever is on that list belongs in front of the servers, not inside them.

what is an mcp gateway: one address, four questions

An MCP gateway is the component an agent host authenticates to before it reaches any tool. Instead of the host holding a separate configuration, credential and connection for every server, it holds one endpoint; the gateway accepts the call, decides what to do with it, and only then forwards it to the server that implements the tool. The phrase names a position in the request path, not a feature list: the gateway is whatever sits between the caller and the server and takes responsibility for the caller's identity.

That position is worth stating precisely, because three different things get called a gateway and only one of them can enforce anything. A reverse proxy forwards bytes and can cap requests per second or per address. A service registry tells a client where a server is. A gateway, in the sense this page uses, reads the MCP message, so it can name the caller, apply a policy keyed to that caller, and write the row that records the call. Parsing the call is the whole distinction; the transport is incidental.

The protocol underneath is deliberately small, which is exactly why the gateway has room to exist. MCP is "an open protocol that enables seamless integration between LLM applications and external data sources and tools" (MCP specification), and its server surface is defined around tools, resources and prompts, with a discovery call and an invocation call for each. None of those primitives defines who may invoke a tool, how frequently, or at what budget. Those are deployment questions, and the gateway is the deployment's answer.

So the shortest honest definition is not "a proxy for MCP traffic". It is "the layer where a deployment answers four questions about a caller before a tool runs": is this caller allowed, which server should serve it, what limits apply, and what is recorded. The next section takes those four apart, because they are what a selection decision actually turns on.

mcp gateways in the request chain: four jobs, four questions

Every mcp gateways implementation can be judged by four jobs, and a component that skips the last two is a proxy that has been renamed.

Authentication decides who is calling. An agent host arrives with a credential — an API key, an OAuth token, a signed assertion — and the gateway resolves it to an identity before any tool runs. Kong's write-up names "centralized identity integration" and OAuth as the first capability of the layer, and it is the job that makes the other three possible: without an identity there is nothing to meter. What a credential has to prove, and how a token is bound to a caller, is worked through on authorization for MCP.

Routing decides which server receives the call. A gateway in front of several servers gives every client one address and dispatches each request to the server that implements the requested tool. Microsoft describes its MCP Gateway as "a reverse proxy and management layer for Model Context Protocol servers" that provides "request routing, authorization and lifecycle management", and LiteLLM describes its own as letting you "use a fixed endpoint for all MCP tools". Routing is the job a load balancer can imitate and the job the next two make non-trivial.

Policy is where rate limits and budgets are enforced — per key, per team, per month. This is the job a byte-forwarding proxy genuinely cannot do: the correct limit is often "this key, this team, this month", which is not expressible in terms of bytes per second. A gateway that parsed the call can hold the counters and refuse the request that would cross them.

Audit writes the record: one row per call, carrying the caller, the route, the tool and the outcome. It is what lets a later question — what did this agent do, what did it cost, was it allowed — be answered from data rather than from memory. As with the tool surface on the wire, the value of the row is that it is written at call time, by the component that saw the call.

Job The question it answers What it needs to run What a deployment loses without it
Authentication Who is calling? A credential and an identity mapping Every caller is anonymous; nothing can be metered per caller
Routing Which server serves this tool? A registry of servers and their tools Each client must know every server address
Policy How much may this caller spend or send? Counters keyed to the identity Limits degrade to per-IP caps and the bill arrives unexplained
Audit What happened, and can we prove it? A row written per call Incidents are reconstructed by guesswork

Those four, and not a ranking of products, are the axis a buying decision should use: read each candidate's documentation and ask which of the four it performs, and where each one is configured.

ai gateway vercel: the model layer, and why it is a different layer

The phrase "AI gateway" usually points at the model-facing layer, not the tool-facing one, and pulling it apart is the fastest way to avoid buying the wrong thing. Vercel describes its AI Gateway as a service to "centralize credentials, log requests, control spend, and fail over across providers" (Vercel AI Gateway); Cloudflare describes the same category as a way to "gain visibility and control over your AI apps" with analytics, logging, caching, rate limiting, request retries and model fallback (Cloudflare AI Gateway). Both route a request to a model. An MCP gateway routes a tool call to a server.

The traffic is different, the credential is different, and the ledger is different. A model gateway holds provider keys and measures tokens; a tool gateway holds server access and measures calls. The two responsibilities do not merge into one object: a model router has no vocabulary for a tools/call message, and a tool gateway has no vocabulary for provider failover. Some products straddle both, but each half is still governed by its own knowledge, and a selection decision has to say which half is being bought.

This is the boundary the head phrase is really asking about. If what you need to control is which model answers and what its tokens cost, the model-facing layer is the product class; if what you need to control is which agent may call which tool, how often and on whose key, the tool-facing layer is. Naming which layer you are buying, before comparing names, removes most of the confusion in this market — and it is why a page about the MCP gateway should say plainly that the AI gateway is not it.

vercel ai gateway models and mcp tool catalogs: two inventories

A model-facing gateway exposes an inventory of models. Vercel's gateway, for example, is browsed as a catalog of provider models and a request names one of them (Vercel AI Gateway). An MCP gateway exposes an inventory of tools: a discovery call returns what the servers offer, and an invocation call names one of them. The two inventories can look alike in a console and mean entirely different things. A model catalog is a routing table over providers a team pays; a tool catalog is an authorization surface over servers a team operates.

That difference decides what "usage" even means, and it is where a shared budget gets confusing. Model usage is measured in tokens and attributed to a provider invoice. Tool usage is measured in calls and attributed to a caller key and a server. A team that tries to hold both in one number without separating the ledgers will watch the two move for unrelated reasons: one long generation and a loop of small tool calls are not the same event even when both are counted as "requests". Keeping the two ledgers distinct is what makes a cost question answerable at all.

For evaluation, the practical test follows directly. When comparing candidates, ask which inventory the product reads and can enumerate. A tool catalog needs knowledge of the servers and their schemas; a model catalog needs knowledge of providers and their models. A product that lists models has told you, by its own documentation, which of the two layers it lives in — which is a fact you can check before spending anything.

gateway ai: three layers behind five names

The phrase "gateway ai" surfaces five overlapping names — AI gateway, LLM gateway, model gateway, MCP gateway and MCP proxy — and the quickest route to disappointment is to assume they are synonyms. They are not. Behind the names are three layers, and the layer decides what the product can enforce.

Name used in the market The layer What flows through it The unit of accounting
AI gateway, LLM gateway, model gateway Model-facing Prompts and completions Tokens, attributed to a provider
MCP gateway Tool-facing MCP messages, including tool calls Calls, attributed to a caller key and a server
MCP proxy Transport Bytes Bytes or connections

A proxy sits at the transport layer, which is why its limits are expressed in transport terms: requests per second, connections, bytes. The moment the correct limit is "this key, this team, this month", the proxy has to be told by a component that parsed the call — and that component is the gateway at the layer above it. The two can sit in the same stack without conflict, but only one of them can answer a question about a caller, and it is not the proxy.

Reading a vendor page against this table is the fastest evaluation there is. Find which layer the product's own documentation describes — prompts, or MCP messages, or bytes — and you have classified it without needing a benchmark. The table also explains why a "best MCP gateway" listicle can mix model routers and tool routers and still look coherent: it is listing names, not layers.

cloudflare mcp gateway and the edge shape

The phrase "cloudflare mcp gateway" points at a deployment question more than a different layer: where does the gateway run? A gateway can live in a team's own network next to the servers, or run as an edge service that terminates the session close to the client. Cloudflare documents both halves of that question — an AI Gateway for model traffic and a guide to building a remote MCP server on the Streamable HTTP transport (Cloudflare Agents). The part that matters for a gateway is the transport, because MCP transports are bindings and not semantics: a transport "defines how messages are framed and delivered ... It does not define what the messages mean" (MCP transports).

That has a direct consequence for where the gateway belongs. If the gateway terminates the session at the edge, then the session state and the counters have to sit somewhere reachable from every edge location, or the limits quietly become per-location and a fleet of clients shares no ceiling. If the gateway runs beside the servers, the edge has less to know and the round trip is longer, but every counter lives in one place. Neither choice is a ranking of products; it is the same trade every stateful service placed near the edge makes, and it decides which of the four jobs are easy and which are hard.

For a selection decision, the question to put to any hosted option is therefore not "is it at the edge" but "where do the counters and the audit rows live, and can every location see them". A gateway whose limits are local to one point of presence is a gateway whose policy job is only half done.

pydantic ai gateway and the single-key model surface

Pydantic AI's gateway is described as "a unified interface for accessing multiple AI providers with a single key, managed through Pydantic Logfire", with requests flowing in each provider's native format (Pydantic AI Gateway). It is a model-facing product. It belongs on this page only because it makes the boundary concrete: "one key for many providers" solves a credential problem in the model layer, and that is not the problem an MCP gateway solves.

The object being consolidated is different in each case. A model gateway consolidates provider keys, so one service holds the credentials a team would otherwise scatter. An MCP gateway consolidates server access, so one credential and one address stand in front of many tool servers, and the caller is metered per key rather than per provider account. Both are worth having, and neither substitutes for the other: holding a single model-provider key tells you nothing about which agent called a tool, and holding a single server credential tells you nothing about which model answered a prompt.

In a stack, the two can coexist cleanly, and the split is healthy precisely because neither layer grows a vocabulary it was not designed for. The mistake to avoid is expecting one product's key model to answer the other layer's questions — a mismatch that only becomes visible when someone asks for a record the product never had a reason to keep.

helicone ai gateway and observability as a gateway job

Helicone's gateway is described as "a unified API for 100+ LLM providers" with "intelligent routing, automatic fallbacks, and complete observability built-in" (Helicone AI Gateway). The phrase "observability built-in" is the useful one, because it names the fourth job from earlier — the audit — approached from the logging side rather than the access side. Observability products arrive at the gateway layer from the row; access products arrive from the limit. Both end up writing the same kind of record.

For a buyer, the practical question is which end the product was built from. A gateway whose origin is logging typically records richly and enforces lightly; one whose origin is access typically enforces richly and may record less. The right answer depends on whether the obligation is "explain the spend" or "prove the call was permitted" — the two are different requirements and they pull the product class in different directions. Naming the obligation first is what makes the comparison tractable.

The neutral vocabulary both ends converge on is the OpenTelemetry generative-AI semantic conventions (OpenTelemetry), which describe the span attributes a tracing stack can consume. A gateway that emits those conventions fits an existing observability stack with less bespoke glue, and that fit is worth weighing alongside the enforcement jobs from the earlier section.

What a gateway enforces, read from ours

The four jobs above are general; here is what one of them costs in practice, read on 2026-10-07 from backend/smartgate/core/plan_entitlements.py and backend/smartgate/core/rate_limiter.py.

  • Two dimensions per tier, not one. Every plan carries a per-key request rate and a per-team ceiling: on the four-tier catalog that is 30 and 30 on the free tier, 300 and 600 on Pro, 600 and 3000 on Teams, 1200 and 9999 on Enterprise. The REST write surface is counted separately (20 / 120 / 300 / 600 per minute). A single "requests per minute" number would hide the dimension that actually bounds a fleet.
  • The ceiling is what a fleet of keys cannot escape. Issuing more keys multiplies the per-key allowance and leaves the team ceiling where it is. That asymmetry is the whole point of the gateway layer: without it, a per-key limit is a limit on one client, not on an organisation.
  • Contracts are data. A team's feature overrides merge over the catalog defaults, and the merged result is cached for a minute — so a negotiated rate is a value in a table rather than a code path, which is what makes it auditable later.
  • A refusal says which limit it hit. The team-surface counter refuses with a retry interval and a scope label, so a client can distinguish "slow down" from "you are over the team ceiling" without guessing. A limiter that returns a bare rejection forces every client to invent its own retry policy — and they invent different ones.

Where SmartGate fits

SmartGate is a tool-facing gateway of the layer this page defines: an MCP-native algorithm gateway for token control, traffic shaping and agent audit. It is the place where the four jobs already exist behind one authenticated surface — one Streamable HTTP endpoint, per-key metering, and an audit row per call — so the decision this page is about becomes operational rather than a build. Seven tools are exposed through that one endpoint (smart_fetch, smart_search, smart_context_gate, smart_dedup, smart_budget_guard, smart_memory and smart_pipe), and each call is counted against the caller's key.

The plan table is about limits rather than features: monthly token caps are 2M, 20M, 100M and 200M+, MCP requests per minute per key are 120, 300, 600 and 1200, audit-log retention is 7, 30, 90 or 180 days, and a team can hold 2, 10, 30 or 9999 keys. Those numbers move with the tier, so the pricing page is the authoritative table and should be read rather than this page.

Whether to run the layer yourself is a scope question, and it is answered by the four jobs rather than by a feature tour. If two of the four are already covered by components you operate and trust, the decision is narrower; if none is, the layer is the first thing to add, ahead of any further tooling.

Where each option is documented

This page defines the layer; it does not rank the products that implement it. Each of the following is described by that vendor's own documentation, and the cluster's vendor page collects the reading: AWS's managed MCP gateway, Microsoft's MCP Gateway, Composio's MCP gateway, Kong's MCP gateway, TrueFoundry's MCP gateway and LiteLLM's MCP gateway. Read them for what each one says it does; read this page for what the layer is, so the two questions stay separate.

How to get started

The first three steps need no purchase at all.

  1. Name the obligation. Decide whether the problem is "explain the spend", "prove the call was permitted", or both. That single line tells you whether the policy job or the audit job is the one to weigh first.
  2. List the callers and the servers. Write down every host that calls a tool and every server that implements one. If the list has one caller and one server, a gateway may be premature; if it has more, one address in front of the servers is the smaller change.
  3. Test the four jobs against what you already run. For each of authentication, routing, policy and audit, ask which existing component performs it and where it is configured. The gaps are the specification for whatever you add.
  4. Classify the candidates by layer. Use the three-layer table above to sort every product you are considering into model-facing, tool-facing or transport. Do not compare across layers.
  5. If the answer is a hosted tool-facing gateway, connect one client to one endpoint and watch a single call end to end — start free with an MCP-speaking client, then read the pricing page once real call volume tells you which tier you need.

Frequently Asked Questions

Is an MCP gateway the same thing as an MCP proxy?

No. A proxy forwards bytes and can cap requests per second or per address. A gateway reads the MCP message, so it can attach an identity, apply a policy keyed to that identity, and write a row for the call. The distinction decides what can be enforced: a proxy cannot limit a caller it cannot name.

Do we need a gateway if we only run one MCP server?

Not necessarily. With one server and a handful of trusted callers, the server itself may be enough. The gateway earns its place when the caller count grows, when more than one server is in play, when a limit or budget has to be enforced per caller, or when someone needs a record of what was called.

Is an AI gateway also an MCP gateway?

Usually not. An AI gateway is model-facing: it consolidates provider credentials and routes prompts to models. An MCP gateway is tool-facing: it consolidates server access and routes tool calls to servers. Some products cover both, but they are different layers with different ledgers, and buying one for the other's job is the common mistake.

Where should the gateway run, at the edge or beside the servers?

It depends on where the counters and audit rows can live. A gateway placed at the edge needs its session state reachable from every location, or its limits become per-location. A gateway beside the servers keeps every counter in one place at the cost of a longer round trip. The deciding question is whether every point of presence can see the same counters.

What should we check first when comparing options?

Which of the four jobs each product performs, and where each one is configured. A product that authenticates and routes but does not meter or record is a proxy with a newer name, and the four-job test classifies it without a benchmark.

Limitations

This page is a framework, not a benchmark. No product is ranked, because the right answer depends on which of the four jobs a team already covers and where its counters can live; the vendor comparisons somewhere else in this cluster are the place to read each option, and this page deliberately does not rank them.

The external descriptions quoted above are the wording of each vendor's own public documentation, read at the linked pages, and they describe scope rather than quality. Where a vendor's documentation does not publish a specific number — a quota, a limit, a price — this page does not supply one; the correct source is always the vendor's own page, linked where the product is named.

The three-layer model is a classification aid, not a claim that any product belongs cleanly to exactly one layer in every configuration. Real products overlap, a single vendor may ship both a model-facing and a tool-facing gateway, and the shared word "gateway" is precisely why the classification has to be argued from documentation rather than assumed from a name.

Because this page carries no code excerpt, and the reason is recorded in the Method note below, it makes no line-numbered or implementation-level claim about any other product. The four jobs describe what to look for; they are not a promise that any specific product performs all four well. The one implementation-level section is our own ("What a gateway enforces, read from ours"): it quotes our limits and counters from our source on a dated pass, and is not a comparison with anyone else's.

Sources

  • The Model Context Protocol specification — modelcontextprotocol.io/specification, for the protocol's own definition and its per-primitive request patterns.
  • The MCP transport binding — specification/2026-07-28/basic/transports, for the statement that a transport is a binding and does not define message semantics.
  • The MCP server tools page — specification/2026-07-28/server/tools, for the discovery and invocation calls a gateway reads and routes.
  • Anthropic's introduction of the protocol — Introducing the Model Context Protocol, for the origin and intent of MCP.
  • Kong's definition of the layer — What is an MCP Gateway?, for the "single, secure entry point" framing and the four-role description this page formalises.
  • Microsoft's reference implementation — MCP Gateway, described by its own documentation as a reverse proxy and management layer for MCP servers.
  • LiteLLM's proxy documentation — MCP Overview, for the "fixed endpoint for all MCP tools" description and per-key and per-team permission model.
  • Cloudflare's model gateway and remote-server guides — AI Gateway and Build a Remote MCP server.
  • Vercel's model gateway — Vercel AI Gateway, for the model-facing scope described in its own words.
  • Pydantic's model gateway — Pydantic AI Gateway, for the single-key, multi-provider description.
  • Helicone's gateway — AI Gateway Overview, for the unified-provider-API description with observability built in.
  • OpenTelemetry's generative-AI conventions — semconv/gen-ai, for the span attributes an observability stack consumes.
  • The demand figures above are this project's own paid measurement: DataForSEO Google Ads, United States, 12-month window, measured 2026-10-03, recorded in this project's search_volume.json and research_brief.md. Product behaviour and plan limits were read from the product at the revision the slice run recorded and re-checked against the live pricing page on 2026-10-03.
  • The enforcement section is our own implementation, read on 2026-10-07 from backend/smartgate/core/plan_entitlements.py and backend/smartgate/core/rate_limiter.py (origin/main): the per-key and per-team values per tier, the merge over catalog defaults, and what a refusal returns.

Method note

This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned none of this page's eight sections (0 abstentions, 8 no-slice verdicts): rule A found no unique symbol in the scanned repository for any of the section keywords, and the remote candidate fallback returned generic helpers that are collisions rather than section-specific evidence. A pinned generic would have given the page the shape of a verified article with none of the substance, so every section but one is written from external, linkable sources, which is the house rule for an unpinned section. The exception is "What a gateway enforces, read from ours" — our own implementation, read on 2026-10-07 from backend/smartgate/core/plan_entitlements.py and backend/smartgate/core/rate_limiter.py, stating the per-key and per-team dimensions and what a refusal returns.

Product claims were read read-only from the product source at the revision the slice run recorded in this project's pipeline_results.json, and the plan figures were re-checked against the live pricing page on 2026-10-03; the demand figures are this project's own measurement. Every external vendor statement quoted above is taken from the URL cited beside it. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.

The slice run for this page recorded 0 of 8 sections pinned, 0 abstention(s) and 8 no-slice verdict(s); BLOCKS is empty because the matcher found no unique symbol for any section rather than section-specific evidence, as the Method note above explains.