Databricks AI Gateway: Platform-Native or a Separate Layer?
The Databricks AI Gateway is now Unity Gateway: a governance and control plane built on Unity Catalog that proxies model traffic, extends permissions to agents and tool calls, and writes usage records into the catalog. That shape answers well when the platform already names your callers and holds your data. It stops answering when callers arrive from outside the platform.
Short answer: The Databricks AI Gateway is now Unity Gateway: a governance and control plane built on Unity Catalog that proxies model traffic, extends permissions to agents and tool calls, and writes usage records into the catalog. That shape answers well when the platform already names your callers and holds your data. It stops answering when callers arrive from outside the platform.
Key takeaways
- The phrase "databricks ai gateway" is a former name, and the rename is the useful clue: here the control plane is the catalog, not a routing box bolted in front of a model endpoint.
- Platform-native governance is cheaper exactly when every caller is already a principal the platform can name and the traffic never leaves its runtime.
- Four things break the moment callers arrive from outside: no principal, no object for the quota, no record home your reviewers can query, and an exit that becomes a migration.
- Decide on five tests — identity, quota, record, boundary, exit — and if you run both shapes, name which control plane owns the single number that refuses a call.
Our own paid measurement puts this page on eight phrases rather than one, and the shape of that
spread is the finding. In DataForSEO Google Ads data for the United States over a 12-month
window, measured 2026-10-01, "ai gateway" carries 2,400 searches a month at difficulty 53,
"databricks ai gateway" and "unity ai gateway" carry 480 each, "azure ai gateway" 210, "litellm ai
gateway" 170, "aws ai gateway" 140, and "truefoundry ai gateway" and "portkey llm gateway" 110
apiece — all recorded in this project's search_volume.json. The list is three questions wearing one
word — category, platform, product — and this page answers the platform one, where the answer changes
the architecture rather than the shopping list.
The head phrase's SERP settles the first question a reader has: it carries an AI Overview whose
citations are the vendor's own documentation, and a related search — "databricks ai gateway vs
unity ai gateway" — asking whether the object changed. It has not. The reference path still reads /ai-gateway/
while the product is documented as Unity Gateway, which is why this page leads with the rename.
Databricks AI gateway: the rename, and what the platform actually governs
Start from the vendor's own sentence, because the architecture argument is inside it. Databricks describes Unity Gateway, formerly AI Gateway, as a governance and control plane built on Unity Catalog that manages access, cost, security and traffic across AI models, agents and tool integrations. Every word that matters in that sentence is a noun from the data platform: governance, catalog, access, cost.
The documented capabilities follow from that placement rather than from a longer feature list: one proxy and one API surface covering both the platform's hosted foundation models and external providers; permissions and auditing extended to runtime interactions between models, agents and tool integrations such as MCP servers; cost attribution and spend caps; rate limits and input/output guardrails; and usage records written into inference tables that live in the catalog.
Read the last item twice: it is the distinguishing one. Most gateways sell you a dashboard and a retention window and stop there: the record is a product artifact, and its distance from your own systems is a service-level agreement. Here the record is a table inside the same permission model that governs the data the application was built on: a spend question and a data-lineage question become two queries against one catalog, run by the team that already owns both.
The consequence cuts the other way too. A control plane inside a catalog inherits the catalog's reach and nothing beyond it. Traffic that originates on a laptop, in another cloud, or against a provider the platform does not proxy never becomes a row the catalog can describe — not because the product is weak, but because the object it governs is not involved. Hold that boundary line: the rest of this page is about where it falls for a given team. The end-to-end decisions an enterprise deployment has to make are set out on enterprise AI gateway architecture, and this page deliberately stops at the platform-versus-standalone question rather than re-arguing that survey.
Unity AI gateway: the catalog is the control plane
Unity Catalog is a permission and lineage model with concrete objects: catalogs, schemas, tables, grants, service principals, groups, and the workspaces that read them. The architectural claim of a platform-native gateway is that model governance can be expressed in those same objects, and the claim is testable. Take each thing a gateway has to do and ask which existing object does the work.
Identity is where this shape is strongest. The caller is a user, a service principal or a group the platform already knows, with an access-control list someone already reviewed. There is no second directory to maintain, no gateway key to map back to a person, and no orphaned credential when a contractor leaves — revoking the principal revokes the traffic. Any gateway that issues its own keys owes you that mapping, and it is the most commonly underestimated integration in the category.
The record is the second, and it is the most underrated. When usage lands in tables inside the catalog, the audit question stops being a product question: which team spent what last month, which call belongs to this job, who could read the prompt at the time — three SQL questions against objects your reviewers can already see. A standalone gateway answers the same three with an export and a join that someone has to build and maintain.
Quota is the third, and it is where the metaphor starts to strain. A rate limit or a spend cap has to attach to something. If that something is a workspace, a principal or an application the platform already models, the policy review is ordinary work in tooling the team already uses. If the limit you need is "these three vendors, these two clouds, this one budget", the platform models it only to the extent that the traffic is inside it. The question for a design review is not whether the gateway has quotas, but what the limit attaches to — and whether that object is the one your compliance reader reports against. For the definition of the category these jobs belong to, and what separates a gateway from the layers beside it, read LLM gateways; this page assumes that line and argues about where the control plane should sit.
Azure AI gateway: one product, two clouds, and what that moves
The same control plane is documented on Azure Databricks, with its own reference page on Microsoft Learn — which separates two ideas that get blended in conversation: the governance model is a property of the platform, and the platform runs on more than one cloud. What moves between clouds is everything around the model object: the identity provider and its tenant boundary, where the traffic is inspected, which provider catalogue the gateway can reach, and which regulatory regime the resulting record falls under.
What does not move is the part that makes the shape platform-native. The principal is still a principal in that cloud's Unity Catalog, the record is still a table there, and the spend question is still a query against a catalog. Two clouds running the same data platform therefore have two native control planes and two native records, and a budget that has to cover both is not a platform-native number. It is a number some layer outside both platforms has to own.
That is the first place a standalone gateway appears in an otherwise entirely platform-native design, and it appears for a structural reason rather than a preference. It also explains a pattern worth recognising in reviews: a team adopts the platform's gateway for everything inside the platform, then discovers that a second, cross-cloud control plane exists on paper but not in practice, because nobody decided which is allowed to refuse a call. The decision is cheap to make early and expensive to retrofit.
Where the constraint is instead that traffic may not leave your own network, the deployment shape is worked through on deploying a gateway inside your own network and not repeated here. The point that matters here is narrower: a private deployment changes who operates the endpoints, not which object the policy attaches to.
AWS AI gateway: the phrase resolves to an assembly, not a product
On AWS the phrase does not lead to a single governed object, and the honest description is an assembly of parts that each do one job. Model access arrives through Bedrock and its provider catalogue; the front door is an API-management or load-balancing layer; identity is IAM and its roles; the record is whatever log sink or table you route the traffic to. Each piece is individually strong.
The architectural consequence is specific. Identity is native — IAM is not a second directory, it is the directory — but the governance object for model traffic is not a catalog object, so the record has no natural home that already knows about the workload that made the call. Whoever assembles the stack has to choose that home and own the join. That is not a criticism of the assembly; it is the bill for assembling, and it has to be paid by a named team rather than by nobody.
The single-number requirement is what turns this from a diagram into an obligation. Whatever the parts, exactly one of them must hold the counter that refuses a call, and it must be the same number the record reports afterwards. An assembly can satisfy that — a shared store behind the front door will do — but the moment nobody owns the store, the effective limit becomes the limit multiplied by the number of copies of the front door. The responsibility split between a conventional edge layer and a model-traffic layer is drawn on AI gateway versus API gateway; what matters here is that an assembled stack is the opposite of a native control plane.
AI gateway: five tests that decide platform-native versus standalone
Strip the branding off and the decision reduces to five questions, each answerable from documentation or an afternoon in a staging environment. They are ordered the way a design review tends to reach them, and the last two decide the split.
| Test | A platform-native control plane answers it when | A standalone layer earns its place when |
|---|---|---|
| Identity | Every caller is already a principal: user, service principal, group | Callers include machines, other clouds or partners with no principal here |
| Quota | The limit attaches to an object the platform already models | The limit spans vendors, clouds or a budget one platform cannot see |
| Record | One row per call is a table your reviewers already query | The record must join systems the platform does not own |
| Boundary | All traffic originates inside the platform runtime | Some traffic starts on a laptop, in a container or in another cloud |
| Exit | Repointing the endpoint is a setting change | The endpoint is part of the host and moving it is a migration |
Read the table as a dependency order rather than a scorecard. Identity comes first because every later row depends on it: without a principal there is nothing to attach a quota to. The record comes next, because a quota that cannot be explained afterwards is a guess with a number on it. Boundary and exit decide between a single native control plane and a native-plus-standalone pair, and they are the rows most often answered from an org chart rather than from the architecture.
The rule we would apply: split governance into a standalone layer when two or more rows land in the right-hand column, and keep it native when only the boundary row does. One row is a seam to monitor — the outside callers still need a number, and someone must decide where it comes from. Two or more means the native control plane describes a minority of the traffic, and a record covering a minority will not survive its first audit. When you do split, split once: two standalone control planes plus a native one is three answers to one question, which is worse than any single wrong one. The naming trap that makes people split unnecessarily — treating a proxy, a router and a gateway as three interchangeable products — is the subject of the proxy, router and gateway naming ledger.
LiteLLM AI gateway: what a standalone control plane actually owns
The open-source control plane is the shape most often chosen for the boundary and exit rows, and it is worth being precise about what it owns. It holds the provider credentials, so callers never carry a provider key. It exposes one request surface, so adding a vendor is configuration rather than application code. It holds its own counter, so a limit can be enforced before the upstream call. It writes its own record, so a call is attributable afterwards. Those four jobs are the category's definition, and this shape genuinely solves them.
Identity is what it does not solve the way a platform does, and this is the cost checklists skip. Its unit of identity is a key it issued, not a principal in your data platform, so the link from key to person to team is a mapping something has to maintain: trivial at ten keys, a small directory service at two hundred, and the thing that quietly blocks the "which team spent what" question six months later. A standalone layer is not wrong here; it is operating a second identity surface.
The record carries a parallel cost. In platform-native form the record lands where the workload description already lives; here it lands in the gateway's own store, and joining it to the jobs, workspaces or products that produced the traffic becomes an export pipeline owned by whoever owns the gateway. That pipeline is also the only way to answer a retention question, because the window is now a setting on someone else's store rather than a table you govern. The selection criterion is therefore about traffic and callers, not features: choose a standalone layer when the traffic is heterogeneous or the callers are not platform principals, and expect to pay in the identity map and the export. Where the thing you are evaluating sells access to a model catalogue rather than governance of your own calls, that is a different component, drawn on the OpenRouter alternative comparison.
TrueFoundry AI gateway: the managed middle and its contract questions
Between "you operate the control plane" and "the platform is the control plane" sits the managed control plane: a vendor runs the counter and the record, and you consume the result. Applied to the five tests, this shape solves the hard part of the counter — a shared number across instances is a property of the service rather than a datastore you build and page for. What it converts into a question is everything that has to be read back out.
Two questions decide it, and both are answered by a contract rather than a benchmark. First: can the record be exported in a machine-readable form, and what exactly does it contain? A dashboard is not a record, so the useful test is whether one call produces one row with a caller, a model, a token count and an outcome, and whether that row can leave the system. Second: what is the retention default, and what happens to the rows when the subscription ends? A managed layer that can refuse a call but cannot return the reason six months later has satisfied the enforcement job and failed the audit one, and no comparison table measures that.
A third question belongs to architecture rather than procurement: what happens when the control plane itself is unavailable. A gateway that fails open during its own outage has no limit at all in exactly the incident when the limit was the point, and a gateway that fails closed has become part of your availability budget. Neither answer is wrong, but the choice has to be made on purpose and written down. Identity mapping is the fourth cost: a caller's platform principal has to be mapped onto whatever unit the managed service bills, across a boundary you do not control.
Portkey LLM gateway: hosted, one hop outside, and the exit test
The hosted control plane is the far end of the standalone axis: traffic terminates at the vendor and is forwarded to the provider from there. The jobs still exist — that is what keeps it a gateway rather than a proxy — but the counter, the record and the credential custody now live somewhere you do not operate, and the prompts traverse a third party. For some organisations that is the correct trade, usually for an operational reason: nobody wants to run a multi-region rate limiter as a side project.
What separates a good instance of this shape from a trap is the exit test, and every standalone layer should face it. Can the same request shape be pointed at a different endpoint by changing one value — a base URL, a header, a project setting? If yes, the gateway is a component you could move and the host is a convenience. If no, the gateway is part of the host's lock-in, which may still be the right call for a small team provided the trade was chosen rather than inherited. Ask it before the first production key is issued: afterwards the question has an organisational answer.
This is also the shape that interacts worst with a platform-native gateway, and the failure mode is predictable. Two systems then describe the same traffic — one holds the credential and the row, the other the permission and the catalog — so the audit answer depends on which system the reviewer happens to ask. If you run both, write down which one owns the number that refuses a call, and make the other a mirror rather than a second opinion.
Where SmartGate fits
Our own product sits on the standalone side of this split, and the honest thing to state is where that side is the right answer. SmartGate is an MCP-native algorithm gateway: one request surface, provider credentials held by the gateway rather than by callers, limits evaluated before the upstream call, and one audit row per call — caller, route, transport, tool, token count, latency, outcome. Enforcement and the record read the same object, which is what makes the counter believable and the audit answerable, and it sees the tool-call surface rather than treating agent traffic as opaque bytes.
The plan table is a set of operational limits rather than a feature list: monthly token caps of 2M, 20M, 100M and 200M+; MCP requests per minute per key of 120, 300, 600 and 1200; audit-log retention of 7, 30, 90 or 180 days; and 2, 10, 30 or effectively unlimited keys per team. Treat /pricing as the authoritative table for all four figures and check it before quoting a retention window to an auditor or a limit to a customer — the numbers here are the ones in force on 2026-09-30, and the pricing page is where a commitment should start.
The rule from the five tests applies to us as much as to anyone. If every caller is a principal inside one data platform, all traffic stays inside it, and the record already lands where your reviewers query — use the platform's own gateway, because a second control plane would add a map and an export for nothing. Our layer earns its place at the boundary: callers that are not principals anywhere, traffic across vendors and clouds, tool calls the platform gateway cannot see, and one budget that has to cover all of it. When both are in play, the number that refuses a call should live on the layer that sees the most traffic, and the other should mirror it.
Frequently Asked Questions
Limitations
This page is an architecture argument, not a product review: no gateway is ranked, and the vendor documentation cited below is where to check whether a capability is available on your cloud, in your region and on your plan. Capability names and packaging in this market change between releases, and the platform in the title has renamed its own control plane once already.
The five tests are a decision aid, not a scoring model. They do not measure whether a deployment is correctly configured, and passing all five in documentation says nothing about whether anyone has tested the refusal path.
Nothing here is a compliance or security guarantee. Where a record lands, who can read it and how long it survives are facts about a specific deployment and its contracts, and the answers live in vendor documentation and in the agreements signed around it.
The plan figures above are operational limits read on one date, not a feature comparison, and they are no substitute for the current table before a commitment. This page carries no code excerpt on purpose; the reason is recorded in the Method note below.
Sources
- Databricks product page for Unity Gateway — databricks.com/product/artificial-intelligence/unity-gateway: the vendor's description of the control plane and its placement on Unity Catalog.
- Databricks documentation, AI governance with Unity Gateway — docs.databricks.com/aws/en/ai-gateway: the reference path, the unified proxy surface, guardrails, rate limits and inference-table usage tracking quoted in section one.
- Databricks engineering notes — Unity Gateway is generally available and expanding agent governance: cost attribution and agent-and-tool governance.
- Microsoft Learn, Azure Databricks — AI governance with Unity Gateway: the same platform object on a second cloud (section three). The permission objects it governs are documented at docs.databricks.com/aws/en/data-governance/unity-catalog.
- AWS documentation — Amazon Bedrock and Amazon API Gateway: the two parts of the assembled stack in section four.
- Standalone, managed and hosted control planes — docs.litellm.ai, docs.truefoundry.com, portkey.ai/docs: each answering the identity, record and retention questions in its own documentation.
- The Model Context Protocol specification — modelcontextprotocol.io/specification: the tool-call surface a gateway has to keep attributable rather than opaque.
- Demand figures and difficulty scores are our own paid measurements — DataForSEO Google Ads, United
States, 12-month window, measured 2026-10-01 — recorded in this project's
search_volume.jsonandresearch_brief.md, with the SERP shape in this project's cached captures. The plan table and retention windows were re-verified against the live pricing page on 2026-09-30, which is authoritative and where a commitment should start.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice
matcher pinned 0 of 8 sections here: rule A found no unique symbol for
any of the eight section phrases. Every one recorded no-slice with an empty local candidate list,
and the remote candidate fallback returned five score-ranked configuration and helper names per
section — the signature of a vocabulary that collides with generic names across a codebase. The run
recorded 0 abstention(s) and 8 miss(es). A pinned generic would have given the page the shape of a verified article with
none of the substance, so every section above is written from sources, which is the house rule for
an unpinned section.
What follows from that honestly: no line-numbered claim, no quoted implementation detail, and no assertion about the internal shape of any product named here. The only figures on the page are this project's own demand measurements and the plan limits re-verified against the live pricing page on 2026-09-30. No batch fingerprints, auction data or internal hosts appear anywhere in this document, so nothing here has to be asserted byte by byte.