SmartGateSmartGate

Kong MCP Gateway: Reusing Your API Gateway for MCP Traffic

A Kong MCP gateway is a Kong Gateway or Konnect deployment with the AI MCP Proxy plugin in front of MCP traffic: MCP requests enter the same plugin chain as your HTTP API traffic and are translated to and from HTTP, so the authentication, rate limiting and logging plugins you already run apply to tool calls.

Short answer: A Kong MCP gateway is a Kong Gateway or Konnect deployment with the AI MCP Proxy plugin in front of MCP traffic: MCP requests enter the same plugin chain as your HTTP API traffic and are translated to and from HTTP, so the authentication, rate limiting and logging plugins you already run apply to tool calls. Kong's public documentation describes both a mode that fronts an existing MCP server and a mode that converts REST endpoints into MCP tools. For a team already operating Kong, the question is not whether the gateway can carry MCP traffic, but which MCP traffic should cross that boundary.

Key takeaways

  • Reusing a deployed API gateway for MCP is an operations decision first: the identity, policy and logging plugins are already there, and MCP is one more protocol on the same control plane.
  • Kong's documented MCP surface has two useful shapes for that decision — proxy an MCP server you already run, or convert a REST API you already expose into MCP tools — and the second is the one that makes an existing gateway pay off fastest.
  • A model gateway and an MCP gateway sit in different layers: routing model calls is not the same as authorising which tool a caller may invoke, and a reuse plan has to name both explicitly.
  • The counter-triggers are real: long-lived sessions, protocol features your gateway does not model, a plugin edition that does not include MCP, or a callers-versus-gateway trust gap all argue for a purpose-built MCP gateway instead of extending the one you have.
  • Write down the one audit row you must be able to produce for a tool call, then check whether your existing gateway can write it at the boundary — that single test decides most reuse cases.

kong mcp gateway: what the Kong route actually adds to MCP traffic

Kong is an API gateway, and the MCP question it answers is a boundary question: when an agent calls a tool, where does that call get authenticated, limited and recorded? Kong's public documentation describes an AI MCP Proxy plugin that runs between an MCP client and an MCP server and translates between MCP and HTTP. Three things about how it is described matter for a reuse decision.

First, it is described as a regular gateway plugin, registered on a Service or Route, not as a separate product you deploy beside the gateway. That is what makes "we already run Kong" relevant: the MCP endpoint inherits the plugin chain the rest of the traffic already uses, so Key Auth or OpenID Connect, rate limiting and logging plugins are the same objects, configured the same way.

Second, the documentation splits the work into two shapes. In one shape Kong fronts an MCP server you already operate and proxies its traffic, which is the right answer when a custom MCP server exists and only needs an authenticated, observable entry point. In the other shape Kong converts REST endpoints you already expose into MCP tools, so an existing HTTP API becomes callable by an agent without a second server to run. Kong also documents aggregating tools from multiple conversion configurations into one served MCP server, which is the shape a platform team reaches for when several teams own tools separately.

Third, the same vendor ships an MCP server of its own for Konnect, so an assistant can query gateway entities and analytics. Read together, these are not five products; they are the ordinary API-gateway pattern — a data plane that speaks one more protocol, a control plane that configures it, and a tool surface for operating it.

What that adds to MCP traffic is concrete but bounded. It adds a shared front door, shared credentials and a shared record. It does not add tool semantics: the gateway decides who may call and what gets logged, while the MCP server still owns what the tool does. Teams that blur those two layers expect a gateway to fix a badly described tool, and it cannot. If the moving parts of the category are still unclear, the definition of an MCP gateway is the page to read first; the rest of this page assumes that boundary and asks whether your existing gateway is the right place to hold it.

mcp gateway aws: the same decision inside one cloud

AWS frames the same problem from the opposite direction: instead of "we already run a gateway, can it speak MCP", it is "we already run our APIs in one cloud, what turns them into tools". AWS publishes prescriptive guidance on MCP deployment patterns, and its documentation describes agent infrastructure whose gateway component converts APIs into MCP tools. The architectural question is identical — reuse the boundary you already control rather than stand up a new one — and the answer differs only in what "already control" means.

For an enterprise that runs most of its traffic inside one cloud, the reuse argument is strong for the same reason it is strong on Kong: identity, network policy and logs already exist in the cloud account, and an MCP endpoint that lives inside that boundary inherits them instead of duplicating them. The weakness is the mirror image: the further your callers, credentials or MCP servers live from that one cloud, the more the reuse becomes a migration rather than a configuration change. A team comparing the two routes should ask which of its MCP servers are inside the same trust boundary as the gateway, and let that answer decide. The vendor-specific split is worked through on the AWS MCP gateway page; the transferable lesson is that "already deployed" is only an advantage when the MCP traffic is also already inside that perimeter.

mlflow ai gateway: a model gateway is not an MCP gateway

MLflow's AI Gateway is documented as a unified interface for LLM providers: one endpoint in front of OpenAI, Anthropic and others, with routing, traffic splitting and fallback. Recent MLflow material extends that story toward agent traffic, including governing which MCP servers an agent can reach and recording tool usage as traces. The distinction worth holding onto is structural rather than vendor-specific: the layer that routes model calls and the layer that authorises tool calls are different workloads.

A model gateway normalises provider APIs, holds provider credentials and balances between models. An MCP gateway parses tool calls, decides which caller may invoke which tool, and writes a record per call. They can be the same process, and increasingly they are sold together, but the questions they answer do not overlap. That matters for reuse because a team can reasonably say "we already have a model gateway" and still have nothing that authorises tool calls — the two sentences sound alike and describe different products.

If the model layer is where your team lives today, the useful next question is whether the tool boundary belongs in it or beside it. A platform that added MCP to an existing serving stack, rather than rebuilding, is TrueFoundry's MCP gateway; reading how that shape is described is a faster way to see the split than reading a feature matrix. Whichever route you take, name the two layers separately in your design document, because a single box labelled "AI gateway" that silently does both is exactly how tool authorisation ends up unowned.

open source llm gateway: what you own when you self-host the front door

Self-hosting an open source LLM gateway buys a front door you control completely: the routing config, the credentials and the logs live in your infrastructure, and no vendor row appears on your bill. For MCP it buys a harder follow-on question, because the front door is only part of what tool traffic needs. You also own the tool registry, the per-caller permission model, the audit record, and the upgrade path when the protocol moves.

That is the honest ledger of self-hosting. Consolidation is the benefit — one process, one config file, one set of API keys for model calls and tool calls — and coupling is the cost, because the same process that routes your completions now also holds the credentials your agents use to reach upstream systems. A gateway that fronts a tool which can write to a production database is a more severe dependency than one that only routes chat requests, and it deserves the same on-call treatment as any other write-capable service.

The practical test is not open source versus managed; it is whether the gateway you operate can express the tool-call boundary without application code. If per-key tool permissions, an audit row per tool call and a hard limit on tool invocations all exist as configuration, self-hosting is comfortable. If any of the three requires a plugin you would have to write and maintain, the maintenance cost is now a permanent line in your budget rather than a one-time integration.

portkey mcp gateway: a registry-and-proxy control plane

Portkey's documentation describes its MCP Gateway as a proxy between MCP clients and MCP servers that handles authentication, access control and logging, with a registry where servers are configured and access is provisioned by workspace and user. The pattern is worth naming because it has become the common shape: a registry of servers plus a proxy that authenticates the caller, injects the right upstream credential and records the call. Once one implementation makes that shape legible, the build question changes from "what would an MCP gateway be" to "which control plane do we extend".

For the reuse decision this matters in a specific way. If your existing API gateway can hold the same two objects — a catalog of upstreams and a policy that maps a caller to the tools it may reach — then the MCP work is configuration, and the registry/proxy pattern is already what your gateway does for HTTP services. If it cannot, a dedicated MCP gateway is not a rival product but a missing control plane, and the honest plan is to add the control plane rather than to bend the API gateway into it.

The adjacent shape is the integration platform: a product whose registry is not gateway services but third-party tools. Composio's MCP gateway is the page for that route. The reason to keep them apart is that a registry of your APIs and a registry of vendors' tools have different owners and different credential lifecycles, and a reuse decision made for one is not automatically right for the other.

ai security gateway: the tool-permission boundary is the product

Strip the marketing from the category and an AI security gateway is one thing: an enforcement point that decides, per caller and per call, whether a tool invocation is allowed and what is written down afterwards. Everything else — prompt inspection, content filters, model routing — is adjacent. For an enterprise that already runs an API gateway, per-tool authorisation is the feature that decides whether reuse is cheap or expensive, because it is the feature that has to exist somewhere and cannot be left to a system prompt.

Two properties make it concrete. The rule must live outside the model, in configuration a security team can review: a caller's permission to invoke a tool is not a sentence in an agent's instructions, and a denial has to be enforced at the boundary rather than suggested to the model. And the resulting record must be per call: caller, tool, arguments, outcome, at a retention you can defend. A gateway that already produces that row for HTTP requests is most of the way there; a gateway that only counts requests is not.

Where the identity stack is the bridge, vendor documentation is the useful reading. Microsoft's MCP gateway is the page for the route that leans on an existing identity and cloud boundary. The governance question — who may call what, and where the single enforcement point sits — is the same question an API gateway was bought to answer, which is precisely why the reuse argument is strongest here and why it deserves the most scrutiny.

cloudflare llm gateway: edge-first AI traffic control

Cloudflare's AI Gateway documentation describes a control plane with caching, rate limiting, logging and provider routing, positioned as an edge service that sits close to callers. Edge-first changes the arithmetic of reuse in two ways. The latency argument is real: policy applied near the caller is cheaper to enforce than a round trip to a central gateway. The trust argument is the opposite: an edge gateway is only a viable MCP boundary if it can hold and inject the credentials the tool calls need, and if the identity you trust at the edge is the identity your tools trust upstream.

The reuse test is therefore narrower at the edge than in a data centre. If your MCP servers are reachable from the edge and your upstream credentials can be stored there under a policy your security team accepts, the existing edge control plane is an efficient MCP boundary. If instead the tool calls have to traverse a private network to reach a database, the edge gateway is a front door that cannot reach the rooms — a useful authentication and rate-limit layer, and the wrong place to enforce tool permissions.

Used carefully, the same control plane can serve both LLM calls and tool calls, which is the consolidation argument again. Used carelessly, it becomes the single process that both holds provider keys and can invoke write-capable tools, with an incident radius nobody sized.

kong ai gateway vs litellm: two ways to front MCP traffic

Kong and LiteLLM are the two clearest answers to the same question, and they differ in where they start. Kong starts from a deployed API gateway: MCP is a plugin on infrastructure you already run, and the MCP endpoint inherits the HTTP control plane. LiteLLM's documentation describes an MCP Gateway as part of a proxy whose primary job is a fixed endpoint for model calls and MCP tools, with permissions managed by key, team or organisation. The comparison is not quality; it is what each assumes you already operate.

Dimension Reuse a deployed API gateway (Kong route) Add an application-layer gateway (LiteLLM route)
What you already run An API gateway, its routes, its plugins and its on-call A model-routing proxy, or nothing yet
Where identity lives The gateway's existing auth plugins and consumers The proxy's keys, teams and organisations
MCP and LLM traffic Same control plane, one more protocol on it One proxy fronting both model and tool calls
Work you take on Configuring and, in some editions, licensing MCP support Operating the proxy, its config and its upgrades
Failure to plan for A gateway trust zone that MCP callers do not share A proxy that becomes the write path to everything

The decision rule that survives contact with a real estate is unromantic: start from the system your team already pages someone for. If that is an API gateway and its documented MCP plugin covers the tools you need, reuse it, because you are buying a new capability on an old operational model. If the thing your team already pages someone for is a model-routing proxy, the equivalent move is there. What neither route should do is add a third thing to operate because a comparison table implied you have to. LiteLLM's MCP gateway carries the other side of this comparison in more detail.

A decision framework: should your existing API gateway carry MCP traffic?

This is the question the rest of the page exists to answer, and it turns on five positive triggers and four counter-triggers. Score your situation against them in order; the first trigger that clearly fires is usually enough to justify a pilot, and any counter-trigger that fires is a reason to stop and design the boundary deliberately.

Test Fires toward reuse when Fires against reuse when
Identity reuse The gateway already authenticates the callers who will call tools MCP callers authenticate somewhere the gateway does not trust
Traffic shape MCP is request-and-response on the same transport the gateway already fronts Sessions are long-lived in a way your gateway's model does not express
Existing upstreams The tools are mostly REST APIs you already expose through the gateway The tools are custom MCP servers with their own handshake and credentials
Policy and record Per-tool rules and a per-call audit row can be written as gateway config Either rule would need a plugin you have to write and maintain
Blast radius An incident at the gateway has a radius you have already sized and accepted Both provider keys and write-capable tools would sit in one process

Run the tests honestly. The most common mistake is to answer the identity test with intent rather than configuration: "the gateway could authenticate the callers" is not the same as "the gateway does". The second most common is to treat the traffic-shape test as a formality, then discover that an MCP session outlives the request the gateway recorded, so the audit row and the real session no longer correspond. Both are configuration questions, and both are cheaper to answer before the pilot than during it.

The output of this exercise should be one page: the triggers that fired, the counter-triggers you accepted, and the boundary that falls out of them. If reuse wins, the pilot is small — put one REST API behind the gateway's MCP conversion, call one tool from one client, and check that the audit row exists. If a counter-trigger wins, the same page tells you which control plane you are buying, and that is a better brief than a feature comparison.

Where SmartGate fits

SmartGate is an MCP-native algorithm gateway, which makes it a different object from a reused API gateway and a natural complement to one. Where an existing API gateway can carry MCP traffic, reuse it; where the tool-call boundary needs per-key limits, per-team budgets and an audit row that a compliance team can read, SmartGate is the layer built for that job rather than adapted to it.

The mechanism is five capabilities powered by seven algorithm primitives. Research is smart_search plus smart_fetch. Context is smart_dedup plus smart_context_gate. Memory is smart_memory, a team-level store. Control is smart_budget_guard, with check, count and record against a hard ceiling. Pipeline is smart_pipe, which chains research, read and remember into one callable job. The operational limits move with the plan — monthly token caps, MCP requests per minute per key, audit log retention and team key counts all change by tier, and the pricing page is the authoritative table. Read those numbers against the audit-row test above, because retention is the column that answers to a compliance deadline rather than to a preference.

The complementary reading is the useful one. A team that reuses Kong for the front door and needs a harder tool-call boundary can put SmartGate behind it; a team that never had an API gateway can start with the MCP-native layer and skip the reuse question entirely.

A gateway's rate limit as a security control, read from our own rate limiter

Kong's MCP plugin inherits the rate-limiting plugins the gateway already runs. Our own gateway writes that limit differently, and the file is short enough to read whole: backend/smartgate/core/rate_limiter.py, read on 2026-10-08 from the gateway that serves this site. It answers the question this page keeps asking — when a boundary decides to refuse a call, what exactly is it counting?

  • Two dimensions, checked in order. MCP calls pass through check_rate_limit_mcp(key_id, team_id, …) with two ceilings: a per-key limit (per_key_limit) and a team ceiling (team_ceiling). The key's bucket is checked first, then the team's, so one runaway client is capped on its own and a whole team is capped together — the two questions this page's decision framework separates.
  • One window, two buckets. Both are fixed sixty-second windows counted in Redis, one counter per key and one per team, each keyed by a time bucket. The per-key bucket protects a single integration; the team bucket is what stops ten well-behaved keys from adding up to one incident.
  • A refusal says which limit fired and when to retry. An over-limit call returns {"allowed": False, "retry_after": …, "limit_scope": …}, where limit_scope is mcp_key or mcp_team and retry_after carries the seconds left on the window. The caller is told which ceiling it hit and how long to wait, rather than handed a bare failure.
  • The numbers are plan data, not hard-coded. The per-key limit and the team ceiling are supplied from the plan entitlements the caller's key carries (plan_entitlements.py), so raising a team's ceiling is a plan change rather than a code change — the same "is this rule configuration or code" test this page applies to any gateway.
  • The REST path has its own, simpler counter. The same module also exposes a team-scoped limit for the rest surface (limit_scope rest_team), so the limit that applies to a tool call and the limit that applies to a REST call are separate, deliberately.

Read plainly, this limit has a narrow brief: it does not decide whether a call is allowed to do something — that is the tool boundary — it decides how much budget one key or one team may spend before anyone notices. That is exactly the piece an existing API gateway has to reproduce for MCP traffic, and reading our own file is a short way to see its shape: two dimensions, one window, and a refusal that names the ceiling it hit.

How to get started

  1. Write the one audit row you must be able to produce for a tool call — caller, tool, arguments, outcome, retention — and keep it to one line. It is the test every reuse decision below runs against.
  2. Inventory the MCP servers and REST APIs you already have, and mark which ones sit inside the same trust boundary as your existing gateway. That answer eliminates half the framework above.
  3. Run the five triggers and four counter-triggers from the framework, and record which fired. If reuse wins, keep the pilot to one tool and one client.
  4. Put one REST API behind the gateway's MCP conversion and call one tool end to end; confirm the authentication, the per-call limit and the audit row all fired, not just that the call succeeded.
  5. If the boundary needs limits and a record your current gateway cannot write, connect an MCP-speaking client to SmartGate and run the same single tool call through it — start free with no purchase, then read the plan table once real call volume tells you which tier you actually need.

Frequently Asked Questions

Is a Kong MCP gateway a separate product from Kong Gateway?

No. Kong's public documentation describes the MCP capability as a plugin that runs on the gateway, sitting between an MCP client and an MCP server and translating between MCP and HTTP. That is why the reuse question is worth asking at all: the MCP endpoint inherits the plugin chain the rest of your traffic already uses, rather than requiring a second gateway to operate.

Does reusing our API gateway mean our MCP servers are now managed by the gateway team?

Not necessarily, and the distinction matters. The gateway owns the boundary: who may call, which limits apply and what gets logged. The MCP server still owns the tool: what it does, how it is described and whether it succeeds. A reuse plan that does not name both owners usually ends with the tool's behaviour blamed on the gateway and the boundary quietly unowned.

Is an AI security gateway just an API gateway with a new label?

Only where the API gateway already enforces per-tool authorisation and writes a per-call record. The security value of an AI gateway is the tool-level rule and the audit row; a gateway that only counts requests can front MCP traffic without answering the question security teams actually ask, which is which caller invoked which tool and whether it was allowed.

Can we run model calls and MCP tool calls through the same gateway?

Yes, and consolidation is often the point, but it concentrates risk. The same process would hold your provider credentials and be able to invoke write-capable tools, so an incident at that process has a larger radius than a model-only gateway. Decide whether you are comfortable with that radius before the configuration is written, not after.

What is the smallest useful pilot for the reuse route?

One REST API you already expose, converted into one MCP tool, called from one client, with a check that authentication, the per-call limit and the audit row all took effect. That is enough to prove the boundary works and to expose the one thing a pilot is meant to find: a rule you assumed was configuration and discovered was code.

Limitations

This page is a decision framework, not a benchmark, and it ranks no product. Kong's MCP behaviour is described here from Kong's own public documentation and is attributed to it; no version number, price, quota or internal implementation detail is asserted, and a reader should treat the vendor's current documentation as authoritative for anything operational.

The framework's triggers are design questions, not guarantees. They tell a team which boundary to build and where to look first; they do not prove that any particular gateway implements per-tool authorisation or a per-call audit row to the standard a security team will demand. Those properties have to be verified in a pilot against the actual configuration, which is why step four of the walkthrough exists.

The demand figures quoted above are our own paid measurements for the United States over a twelve-month window, recorded in this project's measurement files; they describe how many people search, not how much the topic is worth, and they will age. Because this page carries no code excerpt — the reason is recorded in the Method note below — it makes no line-numbered or implementation-level claim about any product named. Finally, no pricing is quoted here: the SmartGate tiers are summarised in the vendor section and the plan table itself should be read from the live pricing page, not from this page.

Sources

Method note

This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned 0 of 8 sections for this page (0 abstention(s), 8 no-slice verdict(s)): local rule A found no unique symbol in the scanned repository for any of the eight section keywords, and the remote symbol-candidates fallback returned generic helpers for every phrase. A pinned generic name would have given the page the shape of a verified article with none of the substance, so every section above is written from sources. Kong's MCP behaviour is reported from Kong's own public documentation, with URLs in Sources and no version, price or quota asserted; the SmartGate figures are the house numbers and defer to the live pricing page. The section keyword under each heading comes from this project's own paid measurement run. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.

One section is our own implementation rather than a citation. "A gateway's rate limit as a security control, read from our own rate limiter" reads our own rate limiter — backend/smartgate/core/rate_limiter.py, with its call site and plan catalog named in that section — and states the two dimensions the limit counts, the window it counts them over, and the fields a refusal returns. It is the page's first-hand segment: the values are ours and the files are named. It describes our own gateway, not Kong's.