LiteLLM Alternative: A Migration Decision, Not a Shortlist
A LiteLLM alternative is a migration decision, not a shopping list. Five criteria decide it: the protocol surface your clients and agents speak, where provider credentials and virtual keys land, where the call record is readable and for how long, where the limits are enforced, and what the move itself has to carry. Answer those five in writing before you compare any product.
Short answer: A LiteLLM alternative is a migration decision, not a shopping list. Five criteria decide it: the protocol surface your clients and agents speak, where provider credentials and virtual keys land, where the call record is readable and for how long, where the limits are enforced, and what the move itself has to carry. Answer those five in writing before you compare any product.
Key takeaways
- Switch when a criterion fails, not when a rival's feature list looks longer: an unsupported protocol, credentials you cannot defend in a review, a record you cannot query or export, or limits applied after the invoice instead of at the call.
- The migration cost lives in the carry-over, not in the cutover: virtual keys, model aliases, routing rules, budget counters and the audit rows your retention policy already covers.
- Guardrail placement decides how expensive a switch is — policy embedded in the gateway's request path has to be reimplemented, policy behind a service call usually does not.
litellm alternative: five criteria before any shortlist
LiteLLM Proxy is a fair reference point for this layer: an open-source, self-hosted service that puts an OpenAI-compatible endpoint in front of many model providers and adds virtual keys, spend tracking and routing on top of it. Its own documentation is the authority on what a given release does. If you run it — or something shaped like it — you already own the four jobs an LLM gateway performs, and the question is not which product is better but when the copy you own stops fitting and what leaving costs.
Four reasons arrive bundled as one search. Coverage: a provider, protocol or transport you need is not supported the way you need it, most often an agent's tool traffic rather than chat completions. Operations: the service wants a datastore and a cache, and someone has to own upgrades, key rotation and the on-call rota. Control: credentials and records sit somewhere that does not survive a security review or a residency requirement. Enforcement: limits and budgets are reported after the fact instead of applied at the moment of the call.
Those four reasons reduce to five criteria — the same five whether you are moving to a hosted router, a platform's bundled gateway or another self-hosted proxy.
- Protocol surface. Which endpoints exist, which one your callers already speak, and whether a tool call survives the hop unchanged.
- Credential landing point. Where provider keys live, who can read them, whether each team holds its own key, and what compromising the gateway container buys an attacker.
- Record. One row per call, in a store you can query, with enough fields to answer a cost question or an incident review later — under a retention window you can defend.
- Limits and budgets. Per-key request windows, per-team token budgets and model allow-lists, evaluated on the request path rather than in a nightly report.
- Carry-over. What the move itself has to take with it: keys, aliases, routes, counters and history.
Write your own answer to all five before you open a product page. The sections below take each criterion in turn, then turn the fifth one into a checklist.
openrouter alternative: a router you rent or a gateway you run
"OpenRouter alternative" is usually the wrong framing, because the two candidates are not the same layer. A hosted router sells one endpoint and one credit balance across many models: no infrastructure to run, no provider keys to hold, and a model catalogue whose availability and revisions the vendor controls. A gateway you operate occupies the same seat in the request path — also one endpoint, also many providers — but it keeps the credential, the routing configuration and the record inside your own tenancy.
The decision reduces to three questions. Where is the credential? With a hosted router the provider keys are the vendor's and you hold one key. Where is the record? You read usage from the vendor's surface, under the vendor's retention. Who can change routing? With a gateway you run, a model alias is a line in a config file you review and can revert; with a hosted router it is a value in somebody else's catalogue.
That is a half-day of verification rather than an evaluation project:
- Read the hosted option's own documentation for two documented facts — whether you can bring your own provider credentials, and what the service retains about a request and for how long.
- Read its pricing page for the fee model (subscription, credit markup or pass-through) and put that answer next to the per-call cost you already know.
- Send one fixed prompt set to both endpoints and compare the response metadata: which model served the call, the token counts, and any routing notes.
- Check whether per-key limits and team budgets exist there, and whether they are enforced or only reported.
Choose the hosted router when no residency or audit-export requirement applies and nobody wants to operate a stateful service. Choose the gateway you run when the credential or the record must stay in your tenancy, when tool traffic needs its own counter, or when exporting the record is a contractual obligation. The same split, worked through against one named router and one named tool gateway rather than in the abstract, is on the OpenRouter alternative comparison.
llm gateway: what the layer actually decides
Remove the product names and an LLM gateway is four functions and one luxury. The functions are one protocol surface for every caller, one credential store, one record per call and one enforcement point for limits. The luxury is routing: aliases, fallbacks, provider selection.
The protocol surface decides more than it looks. Most callers speak an OpenAI-compatible chat API; some speak a provider's own message format; and a growing share of traffic is a tool call rather than a chat completion, carried over a different transport with its own lifecycle. A gateway that handle the first two well and flattens the third into an opaque proxy hop cannot count it, limit it or record it. If agents brought you to this page, test the tool path first. Whether the box in front of that path is a proxy, a router or a gateway also decides what it is allowed to decide, and the three layers are separated in MCP proxy versus router versus gateway.
The credential store decides what a breach costs. Three questions: can provider keys be read from a secret manager rather than a container's environment; does each calling team get its own key, so revocation is per team; and is the administrative surface behind authentication at all. A gateway whose admin API answers on the network is a way to spend your provider quota, not a control.
The record decides which questions you can answer later: one row per call, in a store you can query, carrying caller, route, transport, tool, token count, latency and outcome. An aggregate dashboard is not a record — you cannot export it under a legal hold.
The enforcement point decides whether limits mean anything. A per-key window and a per-team token budget change behaviour only when they are evaluated before the upstream call, and when the counter that refuses a request is the counter the record writes. How those responsibilities divide between a conventional edge gateway and the model-aware layer is compared row by row in AI gateway vs API gateway; what an enterprise deployment has to cover end to end belongs to enterprise AI gateway architecture.
One test worth running before any purchase conversation: stream a tool call, kill the connection halfway, then ask the gateway what it recorded. The honest answer tells you whether the record is a log line or a control.
open source ai gateway: the licence decides less than you think
An open-source gateway removes a licence fee and adds an operations bill. That trade is often worth making, but only once the bill is written down. A licence says who may run the software; it says nothing about who restores a datastore at 2am, who re-issues keys when a provider rotates one, who reviews an upgrade that changes a routing default, or who answers the security questionnaire.
Before adopting one, get five things verified rather than assumed:
- Which store holds keys and counters. Provider credentials, virtual keys, rate-limit windows and budget accumulators frequently share a datastore. Ask what happens to a team's remaining budget when that store is restored from a backup taken a day earlier.
- How configuration is versioned. Routes, aliases and limits should live in a file that goes through code review, with a rollback path that does not need a UI session at 3am.
- What the record can prove. If audit rows live in the same store the gateway writes freely, say out loud that they are an operational log and not tamper-evident evidence.
- The upgrade path. An upgrade that rewrites a configuration schema is a planned outage; knowing the order of operations beforehand separates a window from an incident.
- Who owns the patch cadence. A dependency with a security advisory needs a named owner.
Then run the acceptance test in a staging tenancy with two providers, one team key and a deliberately small budget: send the smoke suite, confirm the counter moves, restart the process mid-stream and read the record, restore the datastore and confirm the counter survived, and rotate a provider key while traffic is live. The deployment mechanics of doing that inside your own network — the endpoint you expose, the base URL that has to fail loudly, the transport contract — are worked through in the private-cloud deployment walkthrough.
The reason to prefer open source is rarely the licence price. It is that the credential, the record and the routing configuration end up in systems you already know how to operate.
portkey ai gateway: the migration questions to answer first
Portkey AI Gateway is one of the options teams evaluate when they leave a self-hosted proxy, which makes it a useful stand-in for any candidate: the questions are structural, and they are the same for every replacement.
Is the routing configuration a file? An alias map that lives in a repository can be diffed, reviewed and reverted. The same map behind a UI is a change with no review and no rollback. Ask which one you are buying before you point forty applications at it.
Can the record be exported? Not "is there a dashboard", but: can one call be fetched by its own id, and can a date range be exported into your own store. If the answer is a screenshot, the record criterion fails whatever the feature list claims.
Are the limits enforced? A per-key window and a team budget computed after the call are reporting. The question is whether the counter that refuses a request is the counter that bills.
Where do the provider keys live, and who can read them? If the answer is "in the gateway's environment", the security review will ask about the container, the deploy pipeline and the backups.
The judgement is not which gateway has more features. Read the candidate's own documentation for those four answers; express your five most-used routes as configuration in both systems; diff the model each alias resolves to; run your smoke suite through both base URLs; and compare the two records for the same call. If the record is worse, the migration is a downgrade however the configuration reads. All of that is an afternoon in staging.
envoy ai gateway: reuse the proxy you already operate
If your north-south traffic already terminates in Envoy or Envoy Gateway, an AI gateway built on that data plane has one strong argument: it is the proxy you already run. One binary to upgrade, one certificate and key path, one place where authentication is enforced, one configuration review process, and the same operational reflexes your platform team has spent years building.
The trade is real and worth stating plainly. Envoy's extension model is built for HTTP-shaped decisions — authentication, rate limiting by descriptor, routing — and its reference documentation is the authority on that. Model-aware decisions have a different shape: counting tokens for the model that served the call, holding a per-key window for tool traffic, applying a team budget. Verify whether those live in the data plane you already configure or in a second component beside it, because that answer decides whether you consolidated anything or added a proxy.
Four checks, each answerable from the candidate's own documentation and a staging deployment:
- Does a per-call record exist, and does it carry token counts?
- Can a per-key request window be expressed in the configuration your team already reviews?
- Do large request bodies and streamed responses survive the filters you attach?
- When a call is denied, is the denial recorded with its reason?
Choose this shape when the gateway is infrastructure you already staff, and when a single data plane matters more than a bespoke policy language. Choose a standalone gateway when the layer needs to be a product with its own roadmap, UI and export path — and accept that you are then operating two proxies.
vercel ai gateway: when the platform gateway is the whole answer
If the application already runs on the platform that ships the gateway, the bundled option answers the first four criteria in the simplest way available: one endpoint, one key, no second deployment, no second credential store, and usage visible where the rest of the deployment is visible. A gateway you do not run is a gateway that does not page you, and that is a genuine advantage rather than a marketing point.
It is the right answer under a narrow set of conditions. The application runs only on that platform and nothing it talks to needs to leave. No compliance regime requires the record to be retained in your own tenancy. Nobody needs per-key tool-traffic counters or a self-hosted audit export. And the platform's availability window is acceptable as the gateway's, because they are the same window.
What to verify before you commit, from the vendor's documentation and a test deployment: whether the limits that matter to you are per key or per project; how usage is exported and how far back; whether you can hold your own provider credentials; what the request and response retention actually is; and what the migration path looks like if one service later has to run elsewhere. That That last question is the one teams skip, and the one that turns a platform choice into a rewrite: the day a workload leaves the platform, a gateway that only exists there leaves with it.
llm guardrails: the layer that decides how expensive a switch is
Guardrails are the part of the stack teams forget when planning a migration, and the part that most often makes it expensive. The reason is placement rather than capability. A guardrail that runs as a hook inside the gateway's own request path is welded to that gateway: moving means reimplementing the policy in the new product's extension API and re-proving it. A guardrail that runs as a separate service the gateway calls is portable, because the contract is a request with a verdict.
Whatever the product, a guardrail is a control only if four things hold. It is enforced outside the prompt, so a user cannot talk past it. Its blocks are recorded — what was flagged, by which rule, on which call — because an unrecorded block cannot be audited. Its failure mode is chosen rather than accidental: when the policy service is unreachable, does the call proceed or stop, and is that decision written down. And it is tested against a fixed set of known-bad inputs.
Ten inputs make a working test set: an instruction-override attempt, an encoded instruction, personal data pasted into the prompt, a request to call a tool the key may not use, and a prompt asking for the model's own system instruction — each at least once. Send them through the deployed path and check three things: that the call was blocked, that the block appears in the record, and that the upstream request never happened. Then send the same ten through the candidate before you migrate, because a policy suite that only passes on the old system is a policy you are about to lose.
Spend belongs in the same place as content policy. A cap that stops a runaway loop is a guardrail; a chart that shows the loop after it finished is telemetry. The per-key windows and team ceilings that make a cap real are the enforcement criterion again, and how those limits resolve per key and per plan is described in what an MCP gateway records and enforces.
What a migration has to carry
The cutover is a maintenance window. The migration is the inventory that makes the window survivable. Nine things move, and each one can be made explicit before the day.
| What to carry | Why it breaks a cutover | How to make it explicit |
|---|---|---|
| Virtual keys and per-key limits | Applications authenticate with them; every key is a credential to re-issue | Export key → team → limits, re-issue in the same window, retire the old set |
| Model aliases | Callers name a model by alias; a changed alias silently changes behaviour | Freeze the alias list and diff the model each one resolves to on both sides |
| Routing rules and fallbacks | A dropped fallback turns a provider incident into an outage | Replay the routes from a configuration file, not from memory |
| Token budgets and counters | A new gateway starts at zero while the month is half spent | Decide whether counters are re-seeded or reset, and who is told |
| Audit rows under a retention rule | Retention does not travel with you; the old endpoint's copy goes when it is shut down | Export the rows your policy covers before decommissioning anything |
| Client configuration | Base URL, key, timeouts, retries and connection limits are per application | Run the same smoke suite through the new base URL for every caller |
| Evaluation and regression cases | They encode what "working" meant last month | Point the existing suite at the new gateway and diff the pass set |
| Guardrail policy and its test set | Embedded policy has to be reimplemented and re-proved | Re-run the ten known-bad inputs through the new path |
| Runbooks and dashboards | On-call needs to know which surface to look at | Rewrite the runbook pages on-call actually opens |
One row in that table deserves a worked example: a host like OpenClaw changes by repointing a single MCP server entry and re-issuing one key, which is the shape connecting OpenClaw over MCP documents.
Sequence it as inventory, shadow, canary, cutover, retire. Shadow first: send a copy of production traffic to the candidate and compare the records for the same calls — this is the step that finds aliases resolving to a different revision. Canary per application, not per percentage, because the unit of migration is an application's configuration. Keep the old endpoint reachable for a week, and write the rollback trigger down before the window opens rather than during it.
Applying the five criteria here
SmartGate is the gateway in this stack, and the honest way to present it on a migration page is against the same five criteria rather than against a competitor's feature list. The protocol surface is one endpoint with tool traffic as a first-class path, so a tool call is counted, limited and recorded instead of tunneled. The record and the enforcement point are the same object: every call writes an audit row, and the per-key window and the team budget read that row.
The plan decides operational limits rather than features. Monthly token caps across the four tiers are 2M, 20M, 100M and 200M+; requests per minute per key are 120, 300, 600 and 1200; audit-log retention is 7, 30, 90 or 180 days; and a team can hold 2, 10, 30 or 9999 keys. The pricing page is the authoritative table — read those four numbers against your own retention obligation and your own peak traffic before treating anything above as a decision.
On carry-over, be equally plain. Applications move by repointing a base URL and re-issuing a key, and the smoke suite you already run is the acceptance test. Budgets are set per team on the new side, so the current month's counters do not travel: if that matters to a reporting cycle, plan the cutover at a month boundary. And if one of the five criteria fails for you — a provider we do not route, a retention window longer than 180 days, a guardrail that must run in your own process — the honest answer is to keep what you have, or run both gateways until the gap closes.
Frequently Asked Questions
Limitations
This page is a decision framework, not a comparison. No product named here is ranked, tested or benchmarked by us, and nothing in it comes from a load test we ran against those products; where a claim about one appears, it is a pointer to the question its own documentation answers. Treat the five criteria as an inventory you fill in with your own evidence, not as a score out of five.
The criteria are not equally weighted and not always all relevant: a team with no residency requirement does not fail the credential criterion, and a single-runtime deployment does not care that a platform gateway is a platform dependency. That judgement is yours.
The plan figures quoted above were re-verified against the live pricing page on 2026-09-30 and are operational limits rather than a feature list. Caps, per-key rates, retention windows and key counts change with the plan, and a compliance decision should read the current table rather than this page.
This page carries no code excerpt, and that is a finding rather than an omission: the slice matcher pinned zero of eight sections, as the Method note below records. The honest consequences follow — no line-numbered claim about any implementation, and no assertion about the internal shape of any gateway named above.
Sources
- LiteLLM documentation — docs.litellm.ai — and its virtual-key reference at virtual keys, for the shape of a self-hosted proxy with per-key credentials.
- The LiteLLM source repository — github.com/BerriAI/litellm, for licence, release cadence and configuration layout.
- OpenRouter documentation — openrouter.ai/docs/quickstart, for how a hosted multi-model router presents its endpoint, catalogue and usage accounting.
- Portkey's AI gateway documentation — portkey.ai/docs/product/ai-gateway, for the configuration and observability surfaces a migrating team has to read.
- Vercel AI Gateway documentation — vercel.com/docs/ai-gateway, for a platform-bundled gateway, its catalogue and its usage surfaces.
- Envoy Gateway documentation — gateway.envoyproxy.io — and Envoy's external-authorization filter reference at ext_authz filter, for what a data plane enforces from configuration.
- The Model Context Protocol specification — modelcontextprotocol.io/specification, for the tool-call transport a gateway has to pass through and count.
- OWASP Top 10 for Large Language Model Applications — genai.owasp.org/llm-top-10, for the input-side risks a guardrail test set should cover.
- NIST's Generative AI Profile, NIST AI 600-1 — NIST.AI.600-1.pdf, for risk vocabulary that maps onto the record and retention criteria.
- The plan figures above, and the pricing table a compliance decision should read for itself: smartgate.network/pricing, read on 2026-09-30.
- Demand figures on this page are our own measurement (DataForSEO Google Ads, United States, 12-month window, measured 2026-09-30), recorded in this project's keyword and brief files.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of 8 sections for this page (1 abstention(s), 7 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the eight section phrases, because this lane's vocabulary — gateway, router, alternative, guardrails — collides with generic configuration and helper names across a product codebase. A pinned generic name would have given the page the shape of a verified article with none of the substance, so every section above is written from sources.
Product behaviour was read from the product source at the revision the slice run recorded in this
project's pipeline_results.json, read-only, and the plan figures were re-verified against the
live pricing page on 2026-09-30. The section phrase echoed in each heading comes from this
project's own paid measurement run, not from a third-party tool. No code, batch fingerprints,
auction data or internal hosts are transcribed, so there is nothing here that has to be asserted
verbatim.