LLM Gateways: What the Category Is and Where Its Edge Sits
An LLM gateway is the component on the model-call path between an application and the model providers, and it does four jobs: it holds the provider credentials, it exposes one request surface for every provider, it evaluates the caller's limit before the upstream call, and it writes one record per call. The category is defined by those four jobs rather than by a product name.
Short answer: An LLM gateway is the component on the model-call path between an application and the model providers, and it does four jobs: it holds the provider credentials, it exposes one request surface for every provider, it evaluates the caller's limit before the upstream call, and it writes one record per call. The category is defined by those four jobs rather than by a product name. A transport proxy, a model router, an agent framework and an observability platform each cover part of that ground, and none of them covers all four.
Key takeaways
- Four jobs define the category: credential custody, a single request surface, pre-call enforcement, and one queryable record per call. Feature lists do not.
- A component that forwards requests but keeps no counter, knows no caller and exports no row is not a gateway, whatever its marketing page says — and the five tests that settle it take an afternoon to run.
- The layers above and beside are different layers: the orchestration framework decides which call to make, the observability platform reads what happened afterwards, and neither decides whether a given call is allowed.
- Where the component runs (your network, a container on a laptop, a vendor's control plane) does not change which jobs exist. It changes who operates the two that must stay single: the counter and the record.
Our own paid measurement puts the phrase plan for this page on a small spread of volume rather than
one word: in DataForSEO Google Ads data for the United States over a 12-month window, measured
2026-09-30, "llm gateways" carries 1,600 searches a month, "vercel ai gateway" 2,400, "docker mcp
gateway" 880, "open source ai gateway" 480 and "portkey ai gateway" 390, all recorded in this
project's search_volume.json. Those are one component of the stack, searched for by three different
routes: as a category, as a deployment shape, and as two named products. This page defines the
category and states where its edge is, because almost every argument about gateways is really an
argument about which of the four jobs a component owns.
LLM gateways: one component on the model-call path, four jobs to do
The component is synchronous and inline. Every model call from an application, a background job or an agent passes through it, and it answers before the upstream provider does. That position, not any feature list, is what makes the four jobs meaningful — each of them is only possible because the gateway is in the path:
| Job | What it means in practice | What is broken without it |
|---|---|---|
| Credential custody | Provider keys live in the gateway or in the secret store it reads; callers hold gateway-issued keys | Every caller ships with a provider key, and rotation becomes an application release |
| One request surface | Callers speak one request shape; the gateway translates to each provider's own shape | Every provider integration is application code, and adding a model is a sprint |
| Pre-call enforcement | The limit is evaluated before the upstream call, against a counter that every instance shares | A limit checked after the fact is a report, not a control |
| One record per call | Caller, route, provider, model, token counts, latency and outcome in a store you can query | Cost attribution and incident review become a collection project |
Read the table as a definition and the boundary writes itself. A component that holds the credential but not the counter is an authenticated proxy. A component that holds the counter but not the record cannot answer the question the counter raises — who spent the budget. A component that holds both but sees requests as opaque bytes cannot fill either, because it cannot tell a chat completion from a tool call.
That last point is the one most often skipped at selection time. A gateway whose request surface covers chat completions but treats tool calls as an opaque pass-through cannot count them, limit them or record them; the tokens are real, the spend is real, and the row that would make it attributable does not exist. If agents brought you to this page, test the tool path before the chat path, because the chat path is the one every vendor has already solved.
Three things sit outside the definition and are worth naming so they are not confused with it. The gateway does not decide what the application should do next: planning, delegation and loop control belong to the orchestration layer above it, and a gateway that starts holding conversation state has become an application framework with an identity problem. Nor does it create memory or retrieval: those decide what goes into the request, before the request exists. And it is not the model: it routes among models and can translate between their surfaces, but a component that owns a catalogue of models and sells access to them is a different commercial object.
Two more consequences of the definition. First, "single" applies to the counter and the record, not to the deployment: an organisation can run one gateway per environment and several per business unit and still be inside the category, provided the number that refuses a request is the same number the record reports. Second, the decisions an enterprise deployment has to make end to end — from tenancy to the plan gate that refuses a call — are set out on enterprise AI gateway architecture, and this page deliberately stops at the category line rather than re-arguing that survey.
The gateway label: five tests before you count a product as one
Naming in this market is loose, and the word "gateway" is now attached to components that are not one. The definition above turns into five questions that can be answered from documentation, from a staging environment, or from one afternoon of testing with two credentials:
- Where does the provider credential live? If the request arrives carrying the caller's own provider key and leaves carrying it unchanged, nothing in the path holds a secret, and nothing in the path can revoke one.
- Does anything in the path know who called? A per-caller identity in the request is the precondition for attribution, for per-owner limits and for a revocation that names one person. A shared static token makes every downstream number a team number.
- Can a call be refused before the upstream sees it? This is the sharpest test. It requires a counter that every instance reads and writes, and it fails for every component deployed as a per-developer or per-pod sidecar, because each copy holds its own number.
- Is there one row per call, in a store you can query and export? An aggregate dashboard is not a record: it cannot be filtered by caller for a specific week, and it cannot be handed to an auditor. Ask what the export looks like before asking what the dashboard looks like.
- Is the administration surface authenticated? A gateway whose admin endpoint answers on the network without authentication is a way for someone else to spend your provider quota, and the credential it holds makes that expensive rather than merely annoying.
A product that fails questions 1 and 2 is a router or a transport proxy: it moves requests, and which of those three names belongs to which layer — and what each layer is allowed to decide — is a separate argument this page does not re-open. A product that passes 1 and 2 and fails 3 and 4 is observability-shaped: it can tell you what happened and cannot prevent anything. A product that passes 3 and 4 but fails 1 has put the counter next to the credential's owner rather than next to the caller's identity, which is a deployment decision you can usually fix rather than a property of the category.
The tests are also the honest way to answer the question most often asked in the other direction: whether a conventional edge gateway or an API-management product already does the work, and which of the four jobs it does not. That comparison is drawn row by row on the responsibility map against a classic API gateway, so this page does not repeat it; the useful thing to carry there is the test list above, because it converts a product-category argument into four verifiable behaviours.
Open source gateway or managed service: the licence is not a capability axis
The most argued-about axis in this category is the least defining one. Whether the software is open source, whether it is self-hosted or sold as a service, and how large the vendor is say nothing about whether the four jobs exist — they say who operates them.
An open-source gateway typically arrives with two of the four jobs already done: credential custody, because the process holds the provider keys, and the single request surface, because one compatible endpoint is the whole point of the project. The two jobs it hands back to you are the counter and the record, and both of them have the same requirement, which is that they must be single. That is an operations sentence, not a licensing one: a counter that must be one number for a team turns into a datastore you run, a backup you own and a migration you plan. Teams that count only the licence fee and the container image consistently under-budget the second half.
Three operational questions separate a library you can operate from a service you can operate:
- Who upgrades it, and what happens when a provider changes? Model providers move endpoints, deprecate response fields and change rate-limit semantics on their own schedule. A self-hosted gateway concentrates those updates in one place — which is a benefit — but only if upgrading is someone's named job rather than an interruption.
- Where does the counter live when there are several instances? If the answer is a local file or in-process memory, the effective limit is the limit multiplied by the number of instances. If it is a shared store, ask what happens during a partition: a gateway that fails open during a datastore outage has no limit at all during exactly the incident when the limit matters.
- How does the record leave the system? Retention and export are the questions an incident review and an audit ask. The house pattern is a row per call with a retention window attached — here the windows are 7, 30, 90 or 180 days depending on the tier, and the shortest is a deliberate limit rather than an oversight — and the pricing page at /pricing is the authoritative table for that, as it is for every figure on this page.
None of that is an argument against open source. It is the observation that the licence axis and the capability axis are orthogonal, and that mixing them produces the two failure modes this category sees most: a self-hosted deployment with no shared counter, and a managed subscription whose record cannot be exported. If the constraint that decides the shape is instead that nothing leaves your own network, the deployment mechanics belong to deploying a gateway inside your own network and are deliberately not repeated here.
The container-local gateway: the category at the far end of one axis
One widely used shape puts the gateway on a developer's machine: a command-line plugin that starts model-context servers as containers and exposes them to a client through a single local endpoint. Applied to the four jobs, it is a partial member of the category, and the honest way to describe it is "the same four jobs at a scale of one".
Credential custody works. Secrets can be injected at launch time, so the client's configuration file never holds an upstream token, and the credential sits in the process that opens the connection — which is exactly job one. The single request surface works for the developer in front of it: one endpoint, one configuration file, a curated set of servers. Pre-call enforcement does not work, and the reason is arithmetic rather than architectural: a per-developer installation has as many counters as there are workstations, and a per-minute window resets whenever a laptop sleeps. One record per call does not work either, because each installation writes its own log, so reconstructing a week of calls becomes a collection job over machines that are not always on.
The rule that decides the split is short enough to remember. If a decision needs a number that must be one for a team, it cannot be made locally; if it needs the context of one person's session, it should be. A local gateway is the right home for the developer's tool list and the launch-time secret, and the wrong home for a monthly token budget. A two-tier arrangement is what most teams end up with, and it should be designed as two tiers from the start rather than discovered when the first budget overrun cannot be attributed.
There is a second, non-obvious cost to the local shape that belongs in a definition page, because it is a category question rather than a routing one: running a stranger's server on your machine means running their code with your credentials nearby. Isolation is a trust decision, and the container shape makes it better than a child process of an editor — but it does not make the shared counter appear. Which layer a local component belongs to, and when a single proxy is the whole stack, is settled where the names are: the proxy, router and gateway naming ledger.
Portkey's gateway and the managed shape: which of the four jobs moves
The hosted shape is the one that changes the most about ownership. You rent the control plane: requests leave your network, terminate at the vendor and are forwarded to the provider from there. The four jobs still exist — that is what makes it a gateway rather than a proxy — but three of them now live somewhere you do not operate.
The counter moves to the vendor's store. That is usually an improvement, because it is a shared number by construction rather than by your own effort. The questions worth asking are the ones whose answers change behaviour: is the counter one number across regions, or one per region; is the per-minute window evaluated per credential or per account; and what does the gateway do when its own datastore is unavailable — refuse, or fail open. The last one matters more than the first two, and it is rarely on a feature page.
The record moves to the vendor's store, and this is where a hosted gateway has to be tested rather than read about. Ask what a row contains, whether it can be exported in a machine-readable form, what the retention default is, and what happens to those rows when a subscription ends. A gateway that can refuse a call but cannot hand you the reason six months later satisfies job three and fails job four, and no benchmark measures that.
The credential arrangement is the third question, and the design that keeps revocation cheap is the one where the provider key stays in the control plane and callers are identified by their own gateway credential. A hosted gateway that asks each caller for the provider key has moved a network hop into the path without moving the identity, which is the worst of both shapes.
Nothing here is a verdict on any named product, and that is deliberate: the questions above are answered by each vendor's own documentation, and the answers change between releases. What this page can state as a category fact is the consequence of the move — switching gateway is a migration, not a setting, because keys, aliases, routes, counters, audit rows and guardrail policy all have to travel with the traffic. The five criteria that make that decision rather than a shopping list, and what the move itself has to carry, are the subject of the criteria for leaving a gateway you already run.
Vercel's gateway and the platform shape: the endpoint arrives with the deployment
The fourth shape is the one that arrives as a feature. When an application is deployed on a platform that also offers a gateway, the endpoint is provisioned alongside the deployment, the credential is usually the platform's own provider relationship, and the model-call path stays inside the same provider network as the functions that make the call. For a small team this is genuinely less work than any of the three shapes above, and pretending otherwise would be dishonest.
Applied to the four jobs, the platform shape passes the first two cleanly and splits the last two in a way worth naming. Credential custody is solved at the platform level, which means the application never holds a provider key — a real improvement over the direct-integration default. The single request surface is solved too, and it is usually the widest surface available: one request shape across every model the platform lists, with the tool-call path included rather than proxied.
Enforcement and the record are where the definition has to be read carefully. A platform gateway's natural unit of identity is the account or the project, because that is the unit the platform bills. That makes the counter shared, which is the hard part — but it also means the number that refuses a request may be the account's spend cap rather than a per-team budget, and the two behave differently inside a large organisation: one account limit is a single number for everyone on it, while a per-team budget lets one team's overspend be visible without being everyone's problem. Ask which unit the limit is evaluated on before assuming it is the one you report against.
The record then decides the exit. The question to ask of any platform-native gateway is not whether it has a dashboard, but whether the same request shape can be pointed at a different endpoint by changing one value — a base URL, a header, a project setting. If it can, the gateway is a component you could move and the platform is a convenience. If it cannot, the gateway is part of the platform's lock-in, which may still be a good trade for the team that chose it, provided the trade was chosen rather than inherited. The wider boundary is worth keeping in view here as well: a hosted service that sells access to a catalogue of models decides which model answers, not whether a call may happen, and the difference between those two components is drawn on the model-router comparison.
Where SmartGate fits
This page has defined a category, so the honest thing to state about our own product is which of the four jobs it performs and which of the four shapes it is. The gateway is the control layer: it holds the provider credential, exposes one request surface for the callers that reach it, evaluates limits before the upstream call, and writes one audit row per call — caller, route, transport, tool, token count, latency, outcome. The enforcement and the record read the same object rather than two systems that have to be reconciled, which is what makes questions three and four answerable at all.
The plan table is a set of operational limits rather than a feature list, and it is the part most often quoted out of context: monthly token caps of 2M, 20M, 100M and 200M+; requests per minute per key of 120, 300, 600 and 1200; audit-log retention of 7, 30, 90 or 180 days; and 2, 10, 30 or effectively unlimited keys per team. Treat /pricing as the authoritative table for all four figures and check it before committing to a retention window or quoting a limit to a customer; the numbers here are the ones in force on 2026-09-30.
For the rest of the picture, the neighbourhood has a reading order: architecture on the centre page, deployment shapes if nothing may leave your network, the responsibility map if the question is what a conventional edge gateway already covers, the migration criteria if the component in production is a self-hosted proxy you are outgrowing, and the naming ledger if you want the three names separated before choosing. This page is the one to read first only if you are still deciding what the word means.
Frequently Asked Questions
Limitations
This page defines a category and draws its edge; it is not a product review and does not rank vendors. Every named product appears because a measured search phrase or a source citation carries its name, and the sentence attached to each one points at the question its own documentation answers rather than at a verdict we are not in a position to give.
The names in this market are loose, and nothing here can be read as a claim that any particular implementation belongs cleanly to one layer. Several products span two of the four jobs at once, and a component's membership can change with a deployment decision rather than a release.
Nothing on this page is a security guarantee. A definition says which job a component can perform, not that it is configured, tested or reviewed, and each of the four jobs needs its own evidence before it can be relied on.
The plan figures are operational limits read from the pricing page on one date, not a feature comparison, and they change. They are no substitute for the current table before a commitment, and this page carries no code excerpt on purpose — the reason is in the Method note below.
Sources
- The Model Context Protocol specification — modelcontextprotocol.io/specification, for the tool-call surface and the transport forms a gateway has to keep countable rather than opaque.
- LiteLLM documentation — docs.litellm.ai, as an example of the self-hosted shape described above.
- Cloudflare AI Gateway documentation — developers.cloudflare.com/ai-gateway, as one hosted implementation of the managed shape.
- Vercel AI Gateway documentation — vercel.com/docs/ai-gateway, as the platform-native shape, where the endpoint is provisioned with the deployment.
- Portkey documentation — portkey.ai/docs, for the questions a hosted control plane answers in its own words.
- Docker's MCP Gateway documentation and repository — docs.docker.com, as the container-local shape and the isolation motivation behind it.
- OpenRouter documentation — openrouter.ai/docs, as the model-catalogue shape this page separates from a gateway for your own traffic.
- Demand figures quoted in this page are our own measurements: DataForSEO Google Ads, United States,
12-month window, measured 2026-09-30, recorded in this project's
search_volume.jsonandresearch_brief.md. - The plan table and the retention windows: taken from this project's brief, which records them as re-verified against the live /pricing page on 2026-09-30. That page is authoritative and is where a commitment should start.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice
matcher pinned 0 of 6 sections here: rule A found no unique symbol for any of the
six section phrases, all six came back no-slice, and the recorded candidate lists were empty rather
than ambiguous, because the vocabulary of this lane collides with generic configuration and catalog
names across a product codebase. It recorded 0 abstention(s) and 6 miss(es), and a pinned generic would
have given the page the shape of a verified article with none of the substance. The house rule for an
unpinned section is to write it from sources, which is what every section above does.
The four jobs, the five tests and the shape-by-shape reading are therefore a sourced argument rather than a reading of our own code, and the only figures on the page are the demand measurements in this project's own paid run and the plan limits re-verified against the live pricing page on 2026-09-30. No code, batch fingerprints, auction data or internal hosts appear here, so nothing on the page has to be asserted verbatim.