SmartGate

Cloudflare AI Gateway: A Managed Gateway, Dissected

Cloudflare AI Gateway is a managed control plane between an application and its model providers: you point the client at a Cloudflare endpoint, and Cloudflare proxies the call, optionally serves it from cache, applies rate and spend limits before the provider is reached, and logs every request.

Short answer: Cloudflare AI Gateway is a managed control plane between an application and its model providers: you point the client at a Cloudflare endpoint, and Cloudflare proxies the call, optionally serves it from cache, applies rate and spend limits before the provider is reached, and logs every request. It is an account-scoped service rather than a component you deploy, and that decides its tenancy, credential scope and compliance edge.

Key takeaways

  • The product is a control plane Cloudflare operates, not software you install: what you create is a gateway configuration with a name, settings and an address, and the traffic path gains a hop that terminates on Cloudflare's network before the provider is called.
  • Caching is off by default and exact-match on provider, endpoint, model, credential header and the whole request body — so a chat product should expect a low hit rate and a template bot a high one.
  • Rate limits count requests over a fixed or sliding window, spend limits count estimated dollars over a window, and both answer with a 429. Only the second one can fall back to a cheaper model.
  • Four boundaries decide whether the hosted shape is right: where the provider key is stored, whether traffic may leave your network, which jurisdiction the logs may live in, and how long an audit trail must be kept.

This page is written around the phrase people actually type: "cloudflare ai gateway" carries 2,400 searches a month in our own paid measurement of United States Google Ads data over a twelve-month window, taken on 2026-10-01 and recorded in this project's search_volume.json. The neighbouring phrases are smaller — "ai gateway" and "vercel ai gateway" sit at the same 2,400, and the rest of the measured set runs from "docker mcp gateway" at 880 down to "litellm ai gateway" at 170. Those eight phrases are the sections below, and they are one question asked eight ways: what does a hosted gateway actually do for you, and at what point is hosting it wrong? The Cloudflare product is the concrete example throughout, because a discussion of managed gateways that names no product stays abstract exactly where the answers matter.

Cloudflare AI Gateway: the layer it is hosted at

Settle what kind of object the thing is first, because every other property follows from it. This is not a binary that runs beside your application and not a sidecar next to your pod. It is an account-scoped control plane that Cloudflare operates on its own network, and what you create is configuration: a gateway has a name (up to 64 characters), settings, and a URL that traffic reaches it through. One account can hold ten gateways on the free plan and twenty on a paid plan, and a client calling the default gateway before one exists has it created on the first authenticated request.

Two request surfaces matter, and they are shaped differently:

Surface Address shape What it is for
Provider-native gateway.ai.cloudflare.com/v1/<account id>/<gateway id>/<provider> Keep the provider's own API path; change only the base URL in your SDK
Unified REST api.cloudflare.com/client/v4/accounts/<account id>/ai/v1/chat/completions One request shape across providers with a standard Authorization header

The consequence of the hosted placement is the traffic path itself. A call that used to travel straight from your application to a provider now reaches Cloudflare's network first, where the gateway applies whatever is configured — a cache lookup, a rate limit, a spend limit, a guardrail evaluation, a route — and forwards what survives upstream. Token counts and per-call records exist because Cloudflare stands in that path, not because the client kept books.

There is a second half to "which layer", and it is the half that gets skipped. The unit of ownership is the Cloudflare account, not the application, not the team and not the network segment, so everything the account scopes inherits that boundary: which upstream credentials a gateway can reach, which configuration a token may change, and which logs exist at all. The decisions an enterprise deployment has to make end to end, from tenancy to the plan gate that refuses a call, are set out on the enterprise AI gateway architecture survey; this page stays on the one concrete example.

Open source AI gateway or a vendor-hosted gateway: where the credential sits

The comparison most teams run first — open source against hosted — is really a question about two things: who holds the upstream credential, and who holds the caller's credential. Cloudflare's answer to the first is generous. Unified Billing spends prepaid Cloudflare credits against supported providers, with a 5% fee applied to the credits you buy and provider inference pricing passed through without markup. BYOK, the store-your-own-keys route, keeps your provider keys in Cloudflare's Secrets Store and references them from the gateway configuration, so a key no longer travels with every request and rotating one is a dashboard or API change rather than an application release.

The second half is where the hosted shape becomes visible, and it is the sentence a self-hosting decision usually turns on. Cloudflare's AI Gateway API tokens are account-scoped: the AI Gateway Read, Run and Edit permissions cannot be restricted to a single gateway. A token holding Run can send requests through every gateway in the account, including gateways configured with stored provider keys, and the documentation says so plainly and points at two isolation routes instead — separate Cloudflare accounts, or a Worker-side AI Gateway binding, which authenticates inside the account without a header.

Self-hosted gateway Cloudflare AI Gateway
Where the provider key lives Your secret store, reached by a process you run Cloudflare's Secrets Store, referenced by gateway configuration
What the caller credential can reach Whatever your code enforces per caller Every gateway in the account
The isolation unit A deployment or namespace you define A Cloudflare account (or a Worker-side binding)
How a key rotates Your release or configuration process A dashboard or API change

Read that table as a boundary rather than a scoreboard: a hosted plane stores and rotates credentials for you and takes tenancy granularity in exchange, and whether the trade is good depends on how many independent tenants you must hold apart — a number your own architecture supplies. Against a conventional edge gateway the comparison is a different one, drawn row by row on the API gateway comparison.

AI gateway caching and rate limits: what the proxy layer decides

Two mechanisms decide most of what a proxy layer is worth: what it can serve without calling the provider, and what it can refuse before the provider is reached.

Caching is disabled by default, and when enabled it is an exact-match cache. The key is built from the provider, the API path, the model, the provider authentication header and the full request body, hashed with SHA-256 — so changing one message, one tool definition or one sampling parameter produces a separate entry rather than a hit. The cache covers text and image responses, a cacheable request is capped at 25 MB, and the TTL runs from a sixty-second floor to a one-month ceiling; per-request headers opt a call in, skip the cache or set their own TTL, and a response header reports HIT or MISS. Semantic caching appears in the documentation as a plan, not a shipped feature.

Cache key input Where it comes from Effect of changing it
Provider The provider segment of the request A separate cache namespace
Endpoint and model The API path and the model id A separate entry
Provider credential header The upstream Authorization header A separate entry per credential
Request body Every message, tool and parameter A separate entry per exact body

That shape tells you where caching pays. A support bot with a fixed prompt template over a small set of answers is the documented good case; a chat product where every turn is a fresh conversation is the opposite, with a hit rate near zero and a cache carried for nothing.

Rate limiting counts requests over a window, evaluated with a fixed or a sliding technique. With ten requests per ten minutes starting at 12:00, ten requests at 12:09 and ten at 12:11 all succeed under a fixed window, because they land in two windows; the second batch fails under a sliding window, because more than ten requests arrived in the preceding ten minutes. Requests over the limit receive a 429 and are not processed. The gateway-level default applies uniformly to every request that reaches that gateway, while narrower limits belong to the routing nodes described below — the difference between a rule that protects a gateway and one that protects an individual route.

What the category does with these mechanisms, as opposed to this vendor's version of them, is the category definition — assumed here rather than restated.

Docker MCP gateway and the container-local shape: a different counter

The container-local shape is the other end of the same axis. Run a gateway on a developer's machine — the shape a command-line MCP plugin takes, starting model-context servers as containers behind a single local endpoint — and the first two jobs work almost by accident: the launch environment injects the secret, and the developer gets one endpoint and one configuration file. The third job does not work, and the reason is arithmetic rather than architectural. A per-workstation installation holds as many counters as there are workstations, and a per-minute window resets whenever a laptop sleeps.

Cloudflare's gateway is the opposite arrangement. Its counter is account-scoped, shared by every request that reaches it whichever machine or environment sent the call, which is why a per-key or per-team limit means the same thing to everybody. The two shapes are not competing products; they answer different questions. The honest design keeps the developer's tool list and the launch-time secret in the local container, where one person's session context belongs, and the team budget and audit line in the account-scoped plane, where a number that must be one for the team can actually be one.

The vocabulary half of this boundary is deliberately not settled here: which of proxy, router and gateway applies to a component, and what each may decide, is the subject of the proxy, router and gateway naming ledger. What the Cloudflare example adds is narrower — the two shapes are complementary tiers, and that split is what makes a per-developer proxy safe to run beside a shared counter.

Portkey AI gateway and other hosted control planes: the questions that transfer

Every hosted control plane is asked the same three questions sooner or later, and Cloudflare's answers are specific enough to show why they matter.

The first is what a limit actually counts. Beside request-count rate limiting, this product offers spend limits: a dollar budget over a rolling or fixed window, evaluated before the request goes upstream, answering with a 429 once the budget is spent. A rule can be scoped by provider, by model, or by custom metadata such as a user, team or application identifier, and each dimension works in one of two modes — split by value, so every distinct value gets its own bucket, or filter by value, so the rule applies only when the dimension equals one given value. Twenty rules fit on a gateway. Two properties deserve to be read before it is relied on: cost is a best-effort estimate from token counts and published model pricing, so the provider's own invoice remains the authority; and enforcement is eventually consistent, so a burst of concurrent requests can briefly exceed the budget before the counter catches up.

The second question is what happens at the limit. The default is to block, but on a dynamic route the answer can differ: when the primary model's budget is exhausted, the request can be routed to a declared fallback model instead of being refused. That changes the failure mode from a visible error to a cheaper answer, which is a product decision rather than an infrastructure one.

The third question is whether routing policy is configuration or code. Here it is configuration: a dynamic route is a named, versioned flow of nodes — conditional branches over the request body, headers or metadata, probabilistic splits for rollouts, model nodes, rate-limit and budget nodes, and an end node — and the route name replaces the model name in the request. Routes are saved as versions and deployed with rollback, which is what lets a limit or a fallback change without a release. Two constraints matter: dynamic routing expects an authenticated gateway with BYOK-stored provider keys, and it is not available on the unified REST API.

Those three questions transfer: any hosted plane can be graded on what its limits count, what it does when they trip, and whether its routing policy is data you version or code you deploy.

Vercel AI gateway and platform-native gateways: the billing boundary

Platform-native gateways arrive as a feature of something else you already bought, and their distinctive property is commercial rather than technical: the gateway and the application sit on the same provider's network and settle on the same invoice. Cloudflare's version of this is visible in how the product has been converging with Workers AI — the company describes the two as moving toward a single control plane for model access, where observability, logging, caching, security and billing controls come from one place, and the unified REST API is the surface that convergence is implemented on. The same account, the same token, the same bill.

The billing mechanics decide what a procurement review will ask. Workers AI requests through a gateway pick a billing mode at creation time: standard billing charges the Cloudflare account at the end of the cycle, while unified billing deducts from a prepaid credit balance in real time. Credits bought through unified billing carry a 5% fee, and inference from third-party providers is passed through at the provider's own per-token rates with no markup. Analytics are exposed on the dashboard and as GraphQL datasets, with requests, tokens, cost, errors and cached-response share as the dimensions.

Two consequences follow. Cost attribution stops at the account unless the request carries custom metadata a spend rule can split on, and metadata is limited to five entries per request. And the exit is cheap in protocol terms but expensive in procedure terms: the same client shape can usually be pointed elsewhere by changing a base URL, while the budgets, routes, keys and log history do not travel with the traffic. A platform-native gateway is a convenience purchased with configuration that lives in someone else's control plane.

Bifrost LLM gateway and self-hosting: the four boundaries

Self-hosting becomes the right answer when one of four boundaries is crossed, and each one can be read off the Cloudflare example. This section names the boundary and the evidence; the mechanics of building the replacement are the neighbouring page's subject.

Credential placement. With BYOK, provider keys live in Cloudflare's Secrets Store. That beats shipping keys with requests, and it is still a boundary: a key that must never leave your own key management system, or must be issued and revoked by your own process, needs the gateway inside your perimeter. The account-scoped token behaviour sharpens the point — the caller credential's blast radius is the account, so a design with per-tenant credentials cannot be built on token scoping alone.

Network path. Every request terminates on Cloudflare's network before it is forwarded. For most internet-facing products that is the advantage; for a deployment whose model traffic must stay on a private network it is disqualifying, and the question becomes what shape the gateway takes when it runs beside the workload. That design space is covered by deploying a gateway inside your own network.

Jurisdiction. This is the boundary least often checked and hardest to work around. Cloudflare publishes a compatibility matrix for its Data Localization Suite, and the AI Gateway row reads: no support for Geo Key Manager, no support for Regional Services, and the customer metadata boundary marked in progress with a footnote saying that jurisdictional storage restrictions for logs are not supported today. Read plainly, a deployment that must guarantee which region handles its traffic, or stores its logs, cannot get that guarantee from configuration here.

Audit retention. Logging is on by default and the controls are real: one request header excludes an individual call, another keeps the metadata while skipping the stored prompt and response, and logs can be deleted from the dashboard or the API. The retention regime is a moving target, though: the documentation splits customers by when their first gateway was created, with later gateways following the Workers Logs schedule and earlier ones keeping the legacy regime — a hundred thousand stored logs across all gateways on the free plan, ten million per gateway on the paid plan, ten megabytes per log and five hundred writes per second per gateway. Your own retention policy is reached by export: Logpush, restricted to Workers Paid accounts, four jobs per account, metered per million requests beyond its allowance. An auditor asking for a two-year-old record in a store you control has to be answered by a design, not by the gateway.

LiteLLM and the self-hosted exit: a threshold table

Taken together, the boundaries above turn into a short list of triggers. The table is the decision, in the order the checks are cheapest to run.

Trigger What it looks like on a hosted gateway What must be true before self-hosting helps
Credential custody Provider keys sit in Cloudflare's Secrets Store; a Run token reaches every gateway in the account You run a secret store with an issuance and rotation process the gateway process can read
Network path Requests terminate on Cloudflare's network The private-network deployment shape is designed before the migration, not after
Jurisdiction No regional processing and no jurisdictional log storage today Your own region-pinned storage exists and the log pipeline writes to it
Retention and audit Retention depends on when the gateway was created; export runs through Logpush You have a sink with its own lifecycle and accept running the export job
Multi-tenant governance One account-wide credential scope, so isolation means separate accounts Per-tenant credentials and budgets are enforceable in your own component

Two honest notes. First, the scale at which hosting stops being enough is not a traffic number; it is the first row that goes red — a small team with one provider, one account and an ordinary retention requirement can run for years on the hosted shape, and moving early buys nothing but operating cost. Second, self-hosting is not automatically more governed: a process you run holds credentials you must rotate, a counter you must keep single across instances and a log you must back up. The reasons to move are the four boundaries, not the desire to own infrastructure.

What the move has to carry — the keys, routes, budgets, records and policy that travel with the traffic — is the replacement and migration criteria, the page this one hands over to.

Where SmartGate fits

Our own gateway is the control layer this page has been describing without naming: it holds the provider credential, exposes one request surface, evaluates a caller's limits before the upstream call, and writes one audit row per call whose fields are the same ones the limits read. Where the hosted example above keeps limits and records inside a vendor's control plane, here the enforcement point and the record are one object, and the isolation unit is a team rather than an account.

The plan table is the part a decision turns on, so it is worth stating in full: monthly token caps of 2M, 20M, 100M and 200M+; MCP requests per minute per key of 120, 300, 600 and 1200; audit-log retention of 7, 30, 90 or 180 days; and 2, 10, 30 or effectively unlimited keys per team. Those four rows are the answer to the credential, retention and governance questions raised above, and the pricing page is the authoritative table for all of them — check it there before committing to a retention window in a contract.

Frequently Asked Questions

Limitations

This page is an anatomy of one managed gateway, not a review of the category and not a ranking. Every Cloudflare behaviour described here is quoted from the vendor's own documentation read on 2026-10-01 and attributed to it; a hosted product changes between releases, and that documentation is the authority for anything a decision depends on.

The four boundaries are stated as triggers rather than thresholds, because the honest part of the answer is qualitative: there is no request-per-second number at which hosting stops being suitable. Two of them — jurisdiction and retention — go stale fastest, so re-read those at the vendor's page rather than trusting them from here.

Nothing on this page is a compliance, security or availability guarantee. The plan figures are the ones in force on 2026-10-01 and move with the plan; the pricing page is the authoritative table for them. And this page carries no code excerpt on purpose — the reason is in the Method note.

Sources

  • Cloudflare AI Gateway documentation — developers.cloudflare.com/ai-gateway, for what the product is and which surfaces it exposes (read 2026-10-01).
  • Cloudflare AI Gateway feature documentation — features · caching · rate limiting · spend limits, for the cache key and TTL bounds, the HIT/MISS header, the fixed and sliding windows, budget scoping, the eventual-consistency caveat and the fallback option.
  • Cloudflare AI Gateway authentication, BYOK and logging documentation — authenticated gateway · store keys · logging, for account-scoped tokens and the isolation routes, the Secrets Store arrangement, the log fields and the per-request headers.
  • Cloudflare AI Gateway limits and pricing documentation — limits · pricing, for the gateway, log, Logpush and DLP ceilings and the billing rules.
  • Cloudflare Data Localization Suite product compatibility matrix — developers.cloudflare.com/data-localization/compatibility, for the AI Gateway row on regional services, key management and the customer metadata boundary.
  • Cloudflare's engineering blog on unifying Workers AI and AI Gateway — blog.cloudflare.com/workers-ai-gateway-unification, for the single-control-plane direction.
  • Demand figures quoted in this page are our own measurements: DataForSEO Google Ads, United States, twelve-month window, measured 2026-10-01, recorded in this project's search_volume.json and research_brief.md.
  • The plan table: taken from this project's brief, which records it as re-verified against the live /pricing page on 2026-10-01. That page is authoritative and is where a commitment should start.

Method note

This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of 8 sections here: rule A found no unique symbol for any of the eight section phrases, all eight came back no-slice (0 abstention(s), 8 miss(es)), and the candidate lists it did return were generic product names from an unrelated codebase rather than a symbol this page could quote. A pinned generic would have given the page the shape of a verified article with none of the substance, so every section above is written from sources, which is the house rule for an unpinned section.

The anatomy of the hosted gateway is therefore a sourced reading rather than a reading of our own code: every Cloudflare behaviour is attributed to the vendor documentation listed under Sources and read on 2026-10-01, and the only figures on the page are the demand measurements from this project's own paid run and the plan limits re-verified against the live pricing page on the same date. No code, batch fingerprints, auction data or internal hosts appear here, so nothing on the page has to be asserted verbatim.