Bifrost AI Gateway: The Self-Hosted Shape, and When It Pays
A Bifrost AI gateway is the lightweight, self-hosted form of this category: one binary speaking a single OpenAI-compatible API in front of twenty-plus providers, with no separate control plane. Whether that shape fits is an engineering decision, settled by four properties — the topology you can restore, the fallback you have tested, where the provider keys land, and the observability hooks you…
Short answer: A Bifrost AI gateway is the lightweight, self-hosted form of this category: one binary speaking a single OpenAI-compatible API in front of twenty-plus providers, with no separate control plane. Whether that shape fits is an engineering decision, settled by four properties — the topology you can restore, the fallback you have tested, where the provider keys land, and the observability hooks you own.
Key takeaways
- "Lightweight" describes the control plane, not the feature list: one process reads a configuration file and writes to a datastore you already operate.
- Self-hosting buys credential custody, a reviewable routing table and a per-call record inside your own tenancy; it adds a process, its datastore and its upgrade path to your on-call rota.
- The two size questions — when self-hosting pays and when a managed control plane is cheaper — cannot be answered before the first four properties are, and scale is the least decisive of them.
- The four properties are: deployment topology, routing and fallback, key custody, and observability hooks. Write your answer to each in one line before you compare any product.
Our own paid measurement puts this page on a spread of phrases rather than one word: in DataForSEO Google Ads data for the United States over a twelve-month window, measured 2026-10-01, "ai gateway" carries 2,400 searches a month, "vercel ai gateway" 2,400, "docker mcp gateway" 880, "bifrost ai gateway" 390, "bifrost llm gateway" 260, "azure ai gateway" 210, "litellm ai gateway" 170 and "portkey llm gateway" 110 — all recorded in this project's own search-volume file. Two of those phrases name a product and several name a shape, which is the split this page is about: the self-hosted, lightweight gateway is a deployment form factor, and the products are documented instances of it. This page describes the form factor and the decisions it forces.
Bifrost AI Gateway: the lightweight self-hosted shape
Bifrost is an open-source gateway — Apache 2.0, maintained by Maxim, written in Go — delivered as one service you run yourself: a container image, an npm launcher, or a library you embed in your own binary. It presents a single OpenAI-compatible API in front of more than twenty providers, from OpenAI and Anthropic to AWS Bedrock, Google Vertex and Azure, and its own published benchmark reports roughly eleven microseconds of added latency per request at a sustained five thousand requests per second. Those numbers are not a verdict and this page does not treat them as one; they describe the artifact. What matters here is that Bifrost is a well-documented reference instance of a form factor — the lightweight, self-hosted gateway — and that the form factor, not the product, is what a team actually has to decide about.
"Lightweight" is a statement about the control plane, not about the feature list. A heavyweight deployment splits the gateway into a data plane that forwards traffic and a control plane that keeps configuration, credentials, budgets and records in a managed service beside it. The lightweight shape collapses the two: one process reads a configuration file and writes to a datastore you chose. That collapse is the entire trade. In exchange for running the process, its datastore and its upgrades yourself, the credential, the routing table and the per-call record stay inside your tenancy.
The definition of the category is settled elsewhere. What a gateway is, and which four jobs make it one rather than a proxy or a router, belongs to the category definition of LLM gateways, and the end-to-end survey of tenancy, identity, configuration and plan gates belongs to the enterprise AI gateway architecture centre page. This page stays on the form factor and asks the four questions that decide whether a self-hosted instance earns its place: the topology you can operate, the fallback you have tested, where the provider keys land, and which observability hooks you are willing to own. The two size questions come last, because they cannot be answered before the first four are.
Bifrost LLM gateway: the topology a small deployment actually needs
A lightweight gateway still has topology, and topology is where "lightweight" stops being free. The deployment surface documented for this shape is wider than a single container: a reverse-proxied container, a Kubernetes deployment with a Terraform module across AWS, GCP and Azure, hosted container targets, and an air-gapped installation. What matters for the decision is not which of those you pick but which stateful dependencies the pick commits you to.
Three stores appear in a typical deployment, and each is a promise to keep. The configuration store holds provider credentials, virtual keys, routing rules and budgets; it can be a file, or Postgres or MySQL. The log store holds the per-request rows, and it is frequently the same database. A single-process deployment with file-backed configuration and no cache is genuinely small; the moment configuration moves into a database, "small" means one process plus one database that somebody has to back up, patch and restore.
The honest rule for a small team is to pick the topology you can already restore. If your platform group runs Postgres with tested backups, a database-backed configuration costs little and buys hot reload and more than one instance. If nobody owns a database on this path, start with the file and a single replica, and treat the move to a shared store as a named project rather than an accident. Multi-node clustering — a gossip-based peer network, automatic service discovery, leader election and state synchronisation — is the point at which this stops being a gateway and becomes distributed infrastructure. Adopt it because availability demands it, not because the diagram looks better. Where the constraint is instead that nothing may leave your own network, the mechanics are worked through in deploying a gateway inside a private cloud and are not repeated here.
AI gateway: routing and fallback that survive a provider outage
An AI gateway's routing layer is two mechanisms usually described as one. Retries stay inside a provider: when a request fails with a transient error — a server error, a DNS failure, a refused connection — the gateway reuses the same credential and waits using exponential backoff with jitter before trying again. Fallback crosses providers: only after the retries are exhausted does the gateway move to the next provider in the configured chain, and each fallback provider gets its own full retry budget.
The subtle half is credential failure, because it is not the same as a server failure. In the behaviour Bifrost documents, an authentication or billing rejection is classified as a permanent per-key failure: the key is marked dead for that request and the gateway rotates immediately to another key in the pool, with no backoff, because waiting cannot revive a bad credential. A rate limit is a transient per-key failure: the key is rotated but a backoff is still applied, because providers commonly enforce account-level quotas shared across keys, so the replacement may not have fresh capacity until the window slides. When every key is permanently dead the gateway returns its own upstream-credentials-exhausted error rather than passing the raw provider status through, which is a genuinely useful distinction for the caller: it separates "your gateway key is wrong" from "the credentials behind the gateway are wrong".
Two engineering consequences follow for any gateway, not just this one. First, a fallback chain is only real if a human has failed a provider on purpose and watched the chain recover. Second, capability parity decides whether the chain is safe — the same request must be servable by the fallback provider, or the fallback converts an outage into a different and quieter failure. Where the request surface itself is the question — what a conventional edge gateway already covers and what the model-aware layer adds — the responsibility map is drawn in AI gateway versus API gateway.
Key custody: where a self-hosted gateway keeps the provider keys
Key custody is the property most often quoted in favour of self-hosting and least often verified, so it is worth being precise about what the lightweight shape actually changes. Provider credentials live in the gateway's own configuration or in the deployment's secret store, and callers never hold them: applications authenticate to the gateway with a gateway-issued key, and the gateway substitutes the upstream credential as it forwards the request. Bifrost makes that caller key a first-class object — a virtual key with its own identifier prefix — and accepts it through the header forms the major SDKs already send, so an existing OpenAI, Anthropic or Google client can point at the gateway by changing a base URL and nothing else.
Virtual keys are also where governance attaches in this shape. Each one can carry its own model and provider allow-list, its own budget with a reset window, its own token and request rate limits, an association with one team or one customer, and an expiry. The container-local variant of the same idea sits at the far end of the same axis: a command-line gateway that starts tool servers as containers on a developer's machine and exposes them through one local endpoint holds one person's secrets and one person's counters. The llm gateways category definition works through that shape in full, so this page records only the boundary: credential custody and a single request surface survive at a scale of one, and a shared budget does not.
What self-hosting does not do is remove the obligation. Running the process that holds the provider keys makes that process, its configuration store, its backups and its deploy pipeline part of the security review. Three questions decide whether the custody claim is real: can the credentials be read from a secret manager rather than a committed file; is the gateway's own administration surface authenticated; and can one team's key be revoked without rotating anything another team uses. A gateway that answers all three is a control. One that answers none has moved a secret into a container and called it governance.
Gateway observability: the hooks a lightweight deployment still owns
Observability is the property teams assume they inherit with a self-hosted deployment and then discover they have to build. The mechanism is usually present — Prometheus metrics for scraping or push, an OpenTelemetry path for distributed tracing, request logging to the log store, and a plugin interface for custom middleware — and a platform gateway exposes the same hooks through its own tracing integrations. But a hook is not a record. The engineering question is what has to be written down for the gateway to be operable, and the list is the same for every implementation.
One row per call, in a store you can query: caller, route, provider, model, token counts, latency and outcome. That row is what answers a cost question, an incident question and a capacity question, and it is what an audit asks for by name. Where the record lives and how long it is kept are policy decisions made before the deployment rather than after it. The four tiers on our own pricing page attach retention windows of 7, 30, 90 and 180 days, and the shortest is a deliberate limit rather than a rounding error. Three hooks are worth owning even in a small deployment. A trace span around the upstream call tells you whether the latency lived in the gateway or in the provider. A counter that moves when a limit is enforced tells you the enforcement point is live rather than decorative. And a log of denials, with the rule that produced each one, is the only artefact that lets you answer "why was this call refused" six weeks later.
LiteLLM and the scale at which self-hosting stops paying
The honest question about a self-hosted gateway is not whether it is better than the alternatives but whether the operations bill is smaller than theirs at your size. LiteLLM is a useful reference point because it is the other widely deployed shape of the same idea — an open-source, self-hosted proxy that puts an OpenAI-compatible surface in front of many providers — and because the work it implies is representative: a process to run, a database for keys and spend, a cache for routing, and a named person to own upgrades.
Below a certain size, that work is trivially paid for. A single team, a handful of applications and one shared credential pool is a case where a container, a configuration file and a nightly backup are an afternoon of setup, and where the alternative — sending traffic through somebody else's control plane — costs a review, a data-flow diagram and a residency argument that often costs more than the container does. This is the band in which self-hosting is the cheap answer, and it is wider than vendor marketing implies.
The bill changes shape as the deployment grows, in three specific ways. Multiple instances turn the counter and the record into shared state, which means a datastore on the request path and a written decision about what the gateway does when that store is unreachable — refusing calls is safe, failing open is not. Multiple teams turn the configuration into a reviewable artefact, because a routing change one team makes is a change every team serves; a configuration file under version control and a UI-only setting behave very differently at three in the morning. And a compliance obligation turns the record into a deliverable with a retention window, an export format and an owner. None of those is a reason to stop self-hosting; each is a reason to plan the transition instead of discovering it. The migration criteria for a team already running a self-hosted proxy, and what the move itself has to carry, are owned by the criteria for leaving a gateway you already run.
Portkey and the managed shape: when renting beats running
The managed control plane is the honest answer more often than self-hosting advocates admit, and the conditions under which it wins are narrow enough to state. You rent the gateway when traffic has no residency constraint that forces it into your own network, when nobody wants to own a stateful service, and when the unit of identity the vendor's control plane offers — usually an account or a project rather than a per-team budget — matches the way the organisation is actually billed.
Portkey is one documented instance of that shape: a gateway with both a hosted control plane and a self-hostable distribution, which is exactly why it is useful here as the reference. The questions to ask of any managed gateway are not feature questions. Is the limit evaluated per credential or per account, and does that match the number you report against? Is the counter one number across regions, or one per region? What does the gateway do when its own datastore is unavailable — refuse, or fail open? And what does a row contain, can it be exported in a machine-readable form, and what happens to those rows when the subscription ends?
There is a crossover, and it is worth naming rather than pretending it does not exist. Below it the managed service is cheaper, because the operations work is real and the compliance argument is weak. Above it the arithmetic inverts: a self-hosted instance amortises across more traffic and more teams, a residency or audit-export obligation forces the credential and the record into your tenancy, and the per-call fee of a hosted control plane becomes a line item a budget review will question. Where the alternative under discussion is not a control plane but a catalogue — a service that sells access to models rather than governing your own traffic — the two objects are different, and the difference is drawn in the OpenRouter comparison.
The platform gateway: what an Azure AI gateway replaces
The fourth shape is the one that arrives with the platform you already run on. Azure API Management documents a set of gateway capabilities for generative AI — token-limit policies, semantic caching, load balancing across backends and token-metric emission — which means the cloud account you already pay for can supply a policy plane in front of model endpoints. When that is true, a self-hosted lightweight gateway has to justify itself against a service that is already deployed, already patched and already inside your network boundary.
Applied to the four properties, the platform shape is strong on two and conditional on the others. Topology is solved: there is nothing new to run beyond configuration on an existing service, which is a real operational advantage and not a marketing point. Key custody is usually solved at the platform identity layer, so the application never holds a provider credential. The conditions attach to the other two. Observability is only as good as the platform's export: the token metrics and logs it emits are useful for dashboards, but the question is whether one call can be fetched by its own identifier and whether a date range can be exported into your own store, because that is what an auditor asks for. And enforcement depends on where the limit is evaluated — a subscription-level cap behaves differently from a per-team budget inside a large organisation, and the two are not interchangeable even when the numbers match.
The decision rule is short. If the constraint that brought you here is cost control across many teams, or an audit obligation with an export requirement, compare the platform's record and its enforcement unit against your own before assuming the platform is cheaper; it may be cheaper to run and more expensive to satisfy. If the constraint is simply that you do not want to operate a second service, the platform gateway is the correct answer and a self-hosted instance is unnecessary work.
A self-hosting checklist, in the order it matters
The ordering is a dependency rather than a preference: each row makes the next one cheaper to decide.
| Decision | The question to answer in writing | Self-host the gateway when | Use a managed or platform gateway when |
|---|---|---|---|
| Topology | Can you restore the process and every store it writes to, on a bad night? | You already operate the datastore and its backups | Nobody wants to own a stateful service on the model path |
| Routing and fallback | Have you failed a provider deliberately and watched the chain recover? | You need provider-level failover you have tested end to end | The vendor's routing is tested and its catalogue covers you |
| Key custody | Where do provider credentials live, and who can read them? | A security review or a residency rule requires them inside your tenancy | Platform identity already holds them and callers never see one |
| Observability | Can one call be exported by identifier, and for how long? | Per-call export under a retention rule is a contractual duty | A dashboard and a token metric are all your review asks for |
| Scale | What is peak traffic, and how many teams share the credential pool? | A fixed operational cost amortises across teams and traffic | Per-call fees stay below the cost of running the service |
Where SmartGate fits
SmartGate is the gateway we operate, and the honest way to place it on this page is against the same four properties rather than against a rival's feature list. The request surface is one endpoint with the tool path as a first-class citizen, so an agent's tool call is counted, limited and recorded rather than tunnelled through opaquely. The enforcement point and the record are the same object: every call writes an audit row, and the per-key rate limit and the per-team token budget read that row rather than a separate reporting pipeline.
The plan table is a set of operational limits rather than a feature list, and it is the part that most often gets quoted out of context. The four tiers carry monthly token caps of 2M, 20M, 100M and 200M+; per-key MCP request rates of 120, 300, 600 and 1200 per minute; audit-log retention of 7, 30, 90 and 180 days; and a team can hold 2, 10, 30 or 9999 keys. Treat the pricing page as the authoritative table and read those four numbers against your own retention obligation and peak traffic before treating any of them as a decision.
Two boundaries are worth stating plainly, since this page is about self-hosting. If a hard requirement is that the record never leaves your own network, a hosted control plane is the wrong answer whatever its feature list says, and the self-hosted form factor described above is the right one. And if the requirement is a per-call export under a retention window longer than the longest tier, that is a gap to close explicitly rather than assume away.
Frequently Asked Questions
Limitations
This page describes a form factor and the decisions it forces; it is not a benchmark and it ranks no product. Bifrost and the other named projects appear because they are documented instances of the shapes under discussion, and every claim attached to one is a pointer to its own documentation rather than a verdict we are in a position to give. The performance figure quoted in the first section is the vendor's own published benchmark, not a test we ran.
Nothing here is a security guarantee. Saying that a gateway can hold credentials, enforce a limit or write a record describes what the software is capable of; whether it is configured, tested and monitored is a property of your deployment.
The plan figures quoted on this page are operational limits read from our own pricing page on 2026-10-01, not a feature comparison, and they change. Caps, per-key rates, retention windows and key counts move with the plan, and a commitment should read the current table rather than this page.
This page carries no code excerpt, and that is a finding rather than an omission: the slice matcher pinned none of its eight sections, as the Method note below records. The honest consequences follow — no line-numbered claim about any implementation, and no assertion about the internal shape of any gateway named above.
Sources
- Bifrost documentation — docs.getbifrost.ai, covering deployment, virtual keys, retries and fallbacks, semantic caching and clustering.
- The Bifrost repository — github.com/maximhq/bifrost, for licence, packaging and the published performance benchmark.
- LiteLLM documentation — docs.litellm.ai, as the reference for a self-hosted proxy with per-key credentials.
- Docker's MCP Gateway documentation — docs.docker.com, as the container-local shape at the far end of the deployment axis.
- Vercel AI Gateway documentation — vercel.com/docs/ai-gateway, as the platform-provisioned shape.
- Portkey documentation — portkey.ai/docs, for the questions a hosted control plane answers in its own words.
- Azure API Management generative-AI gateway capabilities — learn.microsoft.com, for the platform-supplied policy plane discussed above.
- The Model Context Protocol specification — modelcontextprotocol.io/specification, for the tool-call surface a gateway has to count rather than tunnel.
- Demand figures on this page are our own measurement — DataForSEO Google Ads, United States, twelve-month window, measured 2026-10-01 — recorded in this project's keyword and search-volume files.
- The plan table and retention windows: read from our pricing page on 2026-10-01, which is the authoritative source for all four figures.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of 8 sections here (0 abstention(s), 8 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the eight section phrases, because this lane's vocabulary — gateway, routing, fallback, topology, keys — collides with generic configuration, catalog and helper names across a product codebase. A pinned generic name would have given the page the shape of a verified article with none of the substance, so every section above is written from sources.
Product behaviour and the plan figures were read read-only from the product source and re-verified against the live pricing page on 2026-10-01. The section phrase echoed in each heading comes from this project's own paid measurement run, not from a third-party tool. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.