LiteLLM MCP Gateway: What a Self-Hosted Build Owns
LiteLLM's proxy is an open-source LLM gateway that also speaks MCP: it exposes one fixed endpoint for every configured MCP server and controls MCP access by key, team or organization. Taking that route means you operate the gateway process, its database, its rate limits, its records and its upgrades.
Short answer: LiteLLM's proxy is an open-source LLM gateway that also speaks MCP: it exposes one fixed endpoint for every configured MCP server and controls MCP access by key, team or organization. Taking that route means you operate the gateway process, its database, its rate limits, its records and its upgrades. This page describes those obligations from LiteLLM's and Anthropic's own documentation, and where the managed alternative draws the same line differently.
Key takeaways
- LiteLLM's MCP gateway is a feature of a gateway you run yourself: one endpoint for all MCP tools, with permission management at the key, team and organization level.
- The client side is a single environment variable — this is the part that is genuinely easy, and it is also why the operational weight lands on the server you are now responsible for.
- Four obligations define the self-hosted route: deploy it, issue and cap credentials, keep a record, and own the upgrade cycle.
- The upgrade cycle is not optional hygiene: LiteLLM's own production guide tells you to pin an image tag, run migrations once per upgrade, and never rotate the salt key after models are added.
- Before you choose, write down which of those four obligations your team can actually staff; the answer decides the route more reliably than any feature comparison.
litellm mcp gateway: what the open-source route actually is
LiteLLM describes itself as a way to "Call 100+ LLMs using the OpenAI Input/Output Format", and it ships as two things: a Python SDK and a Proxy Server. The proxy is an LLM gateway you run yourself — it listens on port 4000 and answers OpenAI-shaped requests, so any client that already speaks that format can be pointed at it without a bespoke integration. The MCP gateway is a feature of that proxy, not a separate product: LiteLLM's MCP overview says the proxy "provides an MCP Gateway that allows you to use a fixed endpoint for all MCP tools and control MCP access by Key, Team".
Concretely, the documented surface is small enough to state plainly. The gateway handles list tools,
call tools, prompts and resources, and speaks three MCP transports — Streamable HTTP, SSE, and
standard input/output (stdio) — so a remote server and a locally spawned process can both sit behind
the same endpoint. It also exposes direct REST endpoints, /mcp-rest/tools/list and
/mcp-rest/tools/call, so a tool can be invoked with a plain HTTP request and no model in the loop,
which is what makes it testable before an agent is wired to it. Servers are declared in the proxy's
config.yaml under mcp_servers, or added through the Admin UI when database storage is enabled.
The detail that matters operationally is namespacing. LiteLLM prefixes each tool name with its MCP
server's name, so two servers that both expose a search tool do not collide in the tool list; the
documentation notes that server names must now comply with the MCP spec's SEP-986 naming rule, and
that legacy non-compliant names still only warn today but may be blocked later. That is a small
example of the theme this page is about: the gateway is software with a release cycle, and its
release cycle is yours.
What "open source" does and does not mean here is worth separating honestly: the process is something you obtain and run, the configuration is a file you hold, and the data stores are ones you provision — but that is not by itself lower cost, better security or less work. For the category definition rather than one route through it, read the definition of an MCP gateway and its boundaries first; everything below stays on the self-hosted case.
claude code llm gateway: the client side is the easy part
Searching for a way to point Claude Code at your own gateway lands on Anthropic's own documentation,
and it is a good place to see exactly where the work is. The mechanism is one environment variable:
ANTHROPIC_BASE_URL points Claude Code at the gateway, and on a managed machine an admin can add
allowedProviders with the value customEndpoint in the managed settings file so the machine
refuses to talk to anything else — including Anthropic directly or a developer's own proxy. Anthropic
notes that this pinning requires Claude Code v2.1.285 or later. The same page carries a billing nuance
worth knowing: while a gateway credential is active, requests carry that credential instead of a
developer's subscription login, so the subscription's usage limits do not apply to them.
Anthropic's list of what a gateway provides is the page's most useful part, because it states the bargain plainly: credentials stay server-side while developers hold gateway credentials; usage is attributed by developer or team regardless of which provider served the request; and provider switching changes one server-side configuration rather than every developer machine.
Then the same page states the cost of that bargain in one sentence: "the gateway becomes infrastructure your organization operates." It adds two consequences self-hosters should read twice: Anthropic does not endorse, maintain or audit third-party gateway products; and Claude Code adds capabilities with each release, so a gateway that does not forward them breaks the corresponding features. That is the upgrade obligation stated by the client vendor rather than the gateway vendor, which is why it belongs near the top of this page and not in a maintenance appendix. The client side of a self-hosted route is one variable; the server side is a dependency you now track.
google ai gateway and ibm ai gateway: the platform-managed pattern
Two of the highest-intent phrases in this lane are a cloud vendor's gateway and a platform vendor's gateway, and both point at the same archetype: the gateway lives inside an account you already have with that vendor, and the decisions around it — identity, network boundary, region, billing — come from the platform rather than from you. This page does not review those products or claim what they include or cost; the pattern is what the self-hosted route is compared against.
In that pattern the operator's job shifts from running software to configuring policy. You still decide which credentials exist, which servers a key may reach and what the record keeps — those decisions do not disappear — but you are not provisioning a database, patching a container or running a schema migration. The managed route converts a runbook into a settings screen, and the trade is that the enforcement point, the data store and the update cycle sit inside someone else's boundary and release cadence.
The cluster members that own that shape in detail are the platform-managed gateway inside a cloud account and the gateway that lives inside a productivity tenant. Read them beside this page and the contrast is concrete: they describe what the managed pattern hands you, this page describes what stays yours when you decline it. Neither is a verdict — a team whose identity, logging and billing already live in one cloud is often better served by that cloud's gateway, while a team whose reason for a gateway is to keep provider credentials and records inside its own boundary is usually better served by running one. The honest test is four questions: where the provider key must live, where the record must live, who can change credential policy, and who is paged when the gateway is down. Whichever column answers those first is the column to pick.
lasso mcp gateway and lovable ai gateway: managed security, and the bundled gateway
The other two managed phrases in this lane land on two more shapes. One is MCP security sold as a service — a hosted product that inspects, filters or brokers MCP traffic. The other is a gateway bundled with an application builder: the platform you build inside also proxies the model calls, so the gateway is never a separate purchase. As above, this page reviews neither product.
What unites them is what every managed shape abstracts away, and that list is the same as section one's obligations. A managed MCP security product tends to absorb the process, its data store, its patch cycle and often the inspection rules themselves. A bundled app-builder gateway absorbs deployment completely and most credential management with it, but it also inherits the platform's record and data-handling terms, and it is least likely to let you move the enforcement point later.
The managed-connector route is the clearest middle ground and has its own page here: the managed connector gateway for MCP servers describes a service that holds the connections to many MCP servers for you. That is the strongest version of the managed argument, because the fiddly part is the interface to dozens of servers — and it is equally the clearest statement of what self-hosting buys instead: control of the connection list, at the cost of maintaining it. Two properties survive every managed abstraction and stay yours whichever shape you pick: the access decision (which key, team or agent may reach which server) and the exit (what leaves, where it is stored, how you get it out). A self-hosted build answers both by construction and by obligation; a managed route answers both by contract, and the contract is worth reading first.
mcp gateway architecture: the parts a self-hosted build owns
Strip away the branding and a self-hosted MCP gateway is a small number of parts, each of which has to be operated. The value of writing them down is that "we will self-host it" usually means someone has quietly signed the team up for all of them.
| Part | What it is in a self-hosted build | What the managed route covers instead |
|---|---|---|
| The process | A long-running service you deploy, size and keep healthy | The vendor runs and scales it |
| Data stores | A relational database for keys, teams and spend; a cache for shared rate-limit counters | Provisioned and scaled for you |
| Identity and credentials | A master/admin credential plus per-key, per-team and per-org credentials you issue | Your existing platform identity, usually |
| Quotas and budgets | Per-key request and token limits and a spend ceiling you enforce at the call | A policy screen over the same numbers |
| MCP access control | The list of servers each credential may reach, intersected across team, user and org | The vendor's permission model |
| The record | Request and tool-call rows you store, redact and retain for your own window | The vendor's log, on the vendor's terms |
| The upgrade | A release cycle you track, a migration you run, a client contract you keep current | The vendor's release cycle |
The architecture is not the hard part and it is not novel — every row is ordinary backend work. The hard part is that the rows are interdependent: the process enforces the quotas, reads the credentials from the database and evaluates the MCP access list per request, the record is written by the same path that enforces the limit, and the upgrade touches all of them at once. Because any row you change changes the enforcement point, the managed-versus-self-hosted decision should be made against these seven rows rather than a feature list. The category-level treatment of why the rows exist, and where each sits relative to a plain API gateway, is on the gateway architecture page.
mcp gateway docker: containers, data stores and the upgrade path
The "mcp gateway docker" query is the practical one, and LiteLLM's documentation answers it with an
unusually explicit production guide. The official images are published to ghcr.io/berriai and
mirrored at docker.litellm.ai/berriai; the monolithic image bundles the Prisma toolchain, which is
why it is also the image to use when the proxy talks to Postgres. The guide's own instruction on
tags is worth quoting as a rule rather than a detail: pin a version tag, not latest or a moving
tag, so rollbacks are deterministic. It notes the images are signed and that a non-root variant
exists.
Two data stores appear almost immediately. A PostgreSQL database is required for the proxy's auth and tracking features — it holds keys, teams, users and spend logs. Redis is required once you run more than one instance: it shares rate-limit counters, router state and the response cache across instances, and without it each instance enforces limits independently. The guide also spells out a scaling consequence: the connection pool is per worker, so the maximum replica count is also a ceiling on database connections.
The upgrade path is the part most often underestimated, and the documentation treats it as a named
job: a migrations job applies schema migrations against Postgres and runs once per upgrade, while
proxy instances set DISABLE_SCHEMA_UPDATE=true so they never migrate themselves. That split makes a
rolling upgrade safe, and it is a dependency you now own — a proxy running last month's code against
this month's schema is a failure mode that only exists because you operate both. Two values are
effectively write-once: the master key, which authenticates admin calls and is the Admin UI login
password, and the salt key, which encrypts the stored provider credentials and must not be changed
after models are added because that makes them unreadable.
For Kubernetes, LiteLLM publishes a Helm chart that runs the migrations job automatically and keeps schema updates disabled on the proxy pods, with autoscaling and a Prometheus service monitor available through values. This is where the self-hosted route stops looking like a container and starts looking like a platform, and it is worth checking against the two neighbouring product shapes in this cluster: the API-gateway vendor's MCP surface and the platform vendor's gateway both arrive with much of this layer already built. If your team already operates a database, a cache and a chart-driven deployment, the marginal cost of this row is low; if it does not, that is the row that decides the answer.
Keys, quotas and the record: the middle of the runbook
Between deployment and upgrades sits the day-to-day work, the part that never gets a launch post. LiteLLM's virtual-key model is the mechanism: you issue per-entity credentials and attach to each the things you want enforced. The documented fields include a spend ceiling and a budget duration, per-minute request and token limits, and model access; the same fields exist at the team level, and spend is tracked against the key, the owning user and the team in separate tables. The documentation also records a subtlety for anyone designing policy: inheritance is not uniform, and MCP access is evaluated against the key row itself, so a team-level entitlement acts as a ceiling rather than an automatic grant.
That is the self-hosted version of a decision every gateway makes, and LiteLLM exposes a switch for
it. By default a key with no explicit MCP server list inherits its team's list, so the team is a
default every key falls back to. Setting require_key_mcp_access_defined: true flips that: a key with
an empty list is granted no MCP servers at all, and the team's list caps rather than defines the
grant. The documentation recommends the stricter posture and warns that turning it on before keys
carry their own grants will drop MCP access for keys relying on inheritance — a fair summary of
running a policy change inside a system you own: you get the switch, and you get the migration.
The record is the other half of the middle. The proxy writes spend-log rows, and the documentation shows how to redact sensitive information in them and how to ship them onward — to object storage in batches, to a message topic for a warehouse, with retry and drop rules documented in detail. Here the self-hosted route's advantage and its burden are the same fact: the retention window is whatever you configure, and configuring it is your job. LiteLLM's own product pages present single sign-on, dedicated audit logs and multi-team management as features of its commercial offering rather than of the open-source proxy, which is worth reading carefully — it means the free route's record is the operational spend log, and a formal, externally auditable log with SSO in front of it is a different purchase. Neither fact is a criticism; both are inputs to the decision.
The four obligations, in the order they bite
Everything above reduces to four obligations — what "self-hosted" actually means in practice.
- Deploy it. A process, a database, a cache, secrets, health checks, and enough operational familiarity to tell a proxy problem from a provider problem. The documentation makes this tractable — images, charts and a production checklist exist — but tractable is not the same as free.
- Issue and cap credentials. A master credential for admin work, a salt key you never rotate, and per-key, per-team and per-org credentials carrying the limits and the MCP access lists. This is where the security posture actually lives, and it is a policy you now own rather than a toggle you enable.
- Keep the record. Decide what is written, what is redacted, where it is shipped and how long it is kept. A gateway that enforces limits but keeps no usable record is an expense, not a control.
- Own the upgrade cycle. Pin the version, run the migration, watch the client contract, and read the release notes. Anthropic's warning that a stale gateway breaks client features is the clearest statement of why this is ongoing work.
The order is a dependency, not a preference. Deployment is the prerequisite for everything; the credential model determines what the record can attribute; the record is what tells you whether the upgrade changed anything. A team that inverts the order — an upgrade process with no record to check it against — is the common way a self-hosted gateway quietly stops being trustworthy.
How our own gateway counts tokens, and what it took from LiteLLM
The rest of this page describes LiteLLM's proxy from its documentation. This section is the one place
where the relationship is literal, because the gateway that serves this site carries a piece of
LiteLLM inside it: read on 2026-10-08 from
backend/smartgate/modules/budget_guard/algorithm.py.
The file is a deliberate extraction, and its own header records the source of each part — the token
counter from LiteLLM's token_counter.py, the cost function from cost_calculator.py, and the price
data from get_model_cost_map.py. The parts taken are kept as they were: the counting core is the
same line for line, including the rule that a message costs three tokens plus its text, that a name
field adds one, and that every reply is primed with three more. Model-name normalisation is also
LiteLLM's idea, trimmed to what this gateway receives: a missing model falls back to a default
encoding, the GPT-4o and DeepSeek families are routed to the larger vocabulary, and anything else is
resolved through the tokenizer's own model table with the default as its safety net. Prices are read
from one local table shipped beside the file, and the cost is the prompt and completion counts
multiplied by that table's per-token rates.
The modifications are the honest part of the relationship, and there are three. Provider-specific pricing was removed and replaced by the single table lookup, so the code no longer branches on which vendor a model belongs to — which is exactly right for a gateway that is not reselling model access. The model-name fixer was reduced to the input shapes this gateway actually sees. And the tool, image and multi-modal counting branches were dropped, because this gateway counts text and chat messages rather than arbitrary payloads. Everything the page above warns about — the upgrade cycle, the data store, the runbook — is what self-hosting a gateway means; this section is the proof that the counting layer on our side is not a reimplementation of LiteLLM's, but a copy with a short, named list of differences.
Where SmartGate fits
SmartGate is on the managed side of this page's table, and it is worth being direct about that: it is an MCP-native algorithm gateway you connect to rather than one you install, so the four obligations above arrive as operational limits rather than as runbooks. Those limits are exactly what the plan table carries — monthly token caps of 2M, 20M, 100M and 200M+, MCP requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and a team key ceiling of 2, 10, 30 or effectively unlimited. Read those against the record and quota rows above, because retention is the row that answers to a compliance deadline rather than a preference.
This page is not an argument that managing it is better. The self-hosted route keeps the provider key, the record and the enforcement point inside your boundary, and for some teams that is the whole requirement; a runbook you own is not a defect. What the contrast genuinely changes is which column answers your four questions first — where the key must live, where the record must live, who changes policy, and who is paged at 3am. If the honest answer is "we want the control and we can staff it", the open-source route is a real answer and this page is its runbook outline. If the honest answer is "we want the control and we cannot staff the upgrade cycle", the managed route is not a compromise; it is the same control with a different owner. The pricing page is the authoritative table for the limits quoted above.
How to get started
The first three steps cost nothing and produce the answer, whichever route you then take.
- Write the four questions down and answer them for your team: where the provider key must live, where the record must live, who can change credential policy, and who is paged when the gateway fails. This is a twenty-minute exercise that removes most of the debate.
- Map those answers onto the seven-row architecture table above. Every row you cannot staff is either a row you buy or a row that becomes the reason the project stalls in month three.
- Try the client side first, because it is nearly free: point one client at a gateway you control and confirm the credential, the usage attribution and the tool list behave as documented. If the client side alone is already awkward, the server side will not be easier.
- If the answer is the managed route, the fastest test is to connect an MCP-speaking client and run one real call end to end — start free — then read the pricing page once your real call volume tells you which tier you actually need.
Frequently Asked Questions
Is the LiteLLM MCP gateway a separate product I install on its own?
No. It is a feature of the LiteLLM proxy server — the process you deploy. You configure MCP servers in the proxy configuration, or add them through the admin interface when database storage is enabled, and the gateway exposes them behind one endpoint. There is no second service to run, but also no second service to blame when something is wrong.
Do we have to self-host to use an MCP gateway at all?
No, and that is the choice this page is about. Managed gateways exist in several shapes: inside a cloud account, inside a productivity tenant, as a hosted connector service, or bundled with an application platform. Self-hosting is one answer to a control question, not a requirement of the technology.
What is the minimum infrastructure for a self-hosted gateway?
The proxy process starts with very little, but its auth and tracking features need a relational database, and more than one instance needs a shared cache so rate limits and routing state are not per-instance. That database and cache are the real minimum — and the parts people forget when they estimate the effort.
Who updates the gateway when the clients it serves add features?
You do, and it is ongoing rather than occasional. Anthropic's documentation states plainly that a gateway which does not forward new client capabilities breaks the corresponding features, and that the gateway product has to be kept updated as the client evolves. Pinning a version makes rollbacks deterministic; it does not remove the upgrade.
Limitations
This page describes one route through a category and is not a benchmark or a review. It makes no claim about ranking, quality or total cost of ownership, because those depend on where a team's credentials, records and operational habits already live. The four-obligation framing structures the decision; it is not a scored comparison.
Product behaviour quoted here is taken from the vendors' own documentation and is deliberately narrow. No figure about model support, pricing or performance is asserted for any named product beyond what its documentation states, and the platform-managed products named in the middle sections are described by archetype only — this page does not assert their features, limits or behaviour. Documentation ages, and a detail that was accurate when read may have changed.
Because this page carries no code excerpt — the reason is recorded in the Method note below — it makes no line-numbered or implementation-level claim about any gateway. The SmartGate plan limits quoted are operational numbers that change with the plan and should be read from the current pricing page rather than from this page.
Sources
- LiteLLM's MCP overview — docs.litellm.ai/docs/mcp, for the fixed-endpoint description, the MCP operations, the REST tool endpoints, the supported transports and the server-name namespacing note.
- LiteLLM's MCP permission management — docs.litellm.ai/docs/mcp_control,
for per-key, per-team and per-organization access, the intersection model, and the
require_key_mcp_access_definedrecommendation. - LiteLLM's production deployment guide — docs.litellm.ai/docs/proxy/deploy, for the published images, the pin-a-version instruction, the PostgreSQL and Redis requirements, and the migrations-job split.
- LiteLLM's production checklist — docs.litellm.ai/docs/proxy/prod, for the master key, the salt key's write-once behaviour, Redis as the shared state store, and the health-check and job-role guidance.
- LiteLLM's virtual keys and spend tracking — docs.litellm.ai/docs/proxy/virtual_keys, for budgets, per-minute limits, team and organization ownership, the cost fields and where spend is tracked.
- LiteLLM's logging documentation — docs.litellm.ai/docs/proxy/logging, for spend-log redaction and the object-storage and message-topic export paths.
- LiteLLM's product home — docs.litellm.ai for the "Call 100+ LLMs using the OpenAI Input/Output Format" description, and www.litellm.ai/enterprise for the commercial feature list that includes single sign-on, audit logs and multi-team management.
- Anthropic's Claude Code LLM-gateway documentation — code.claude.com/docs/en/llm-gateway, for the environment variable that points a client at a gateway, the managed-settings pin, the credential and usage-attribution value, and the statement that the gateway becomes infrastructure the organization operates.
- The Model Context Protocol specification — modelcontextprotocol.io/specification, for the transports and the tool interface the gateway exposes.
- Demand figures quoted in this page are our own measurements: DataForSEO Google Ads, location 2840,
desktop, depth 20, measured 2026-10-03, recorded in this project's
search_volume.jsonandresearch_brief.md.
Method note
This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned 0 of 8 sections for this page (0 abstention(s), 8 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the eight section keywords, because the vocabulary of this lane — gateway, proxy, key, architecture — collides with generic helper and type names across a codebase. The nearest candidates the matcher returned were unrelated containers, and a pinned generic name would have given the page the shape of a verified article with none of the substance, so every section above is written from external sources.
Every product claim in this page is read from the vendor's own documentation — LiteLLM at docs.litellm.ai and Claude Code at code.claude.com — and the specific pages are listed under Sources. No model-count, price or performance figure is asserted for any named product beyond what those pages state, the platform-managed products named in the middle sections are described by archetype rather than by feature, and the SmartGate plan figures, where quoted, are cross-referenced to the live pricing page. The section keyword quoted above each heading comes from this project's own paid measurement run, not from a third-party tool. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.