Azure OpenAI Pricing: Managed vs Direct, Two Shapes of Cost
Azure OpenAI pricing does not have the shape of a per-token card, even though it prints per-token rates. A managed deployment is billed in several modes at once — Standard pay-as-you-go by input and output tokens, Provisioned throughput charged as an hourly rate per unit of reserved capacity whether or not requests arrive, and a Batch tier discounted against Global Standard pricing — and each…
Short answer: Azure OpenAI pricing does not have the shape of a per-token card, even though it prints per-token rates. A managed deployment is billed in several modes at once — Standard pay-as-you-go by input and output tokens, Provisioned throughput charged as an hourly rate per unit of reserved capacity whether or not requests arrive, and a Batch tier discounted against Global Standard pricing — and each mode carries its own deployment type, its own regional rate and its own subscription quota. The practical consequence is that the same model's unit price on a cloud vendor cannot be compared line-for-line with the model owner's own card, because the metering boundary, the billing granularity and the bundled infrastructure all differ. Compare deployment shapes before you compare a single figure.
Key takeaways
- A managed deployment is a contract, not a rate. One cloud sells the same model three ways — pay-as-you-go tokens, hourly reserved capacity, and a discounted batch tier — and the three do not convert into each other cleanly.
- Provisioned capacity is billed by the hour regardless of use. You are charged for the units you reserve, not the tokens you consume, and a provisioned deployment cannot be paused: billing stops only when you delete it.
- The unit price is regional and per deployment type. Global, Data Zone and Regional deployments price the same model differently, and quota is granted per region and per subscription.
- Quota is not billing. A deployment can hold throughput quota that never appears as a line item, because quota governs what you are allowed to do, not what you owe.
- An enterprise agreement sits on top of the card. Azure's own pricing page states the figures are estimates whose actual value depends on the agreement, the date and the exchange rate.
- Do this first: write down the deployment type and the billing mode each candidate uses before comparing any two numbers, then compare only candidates that bill the same way.
Most confusion about azure openai pricing comes from reading a cloud vendor's page as if it were the model owner's card. They look similar — both quote a rate per million tokens for input, cached input and output — but they are different instruments. The model owner sells inference directly; the cloud vendor sells a managed resource that contains inference plus capacity, quota, a region and a contract. The rate table is the entry point, not the product.
This page is one member of a cluster about what AI costs, and it owns exactly one slice of that question: the difference in shape between a managed or resold price and a direct one. The cluster's centre is the general taxonomy of billing mechanisms; reading the columns of a single card, the arithmetic of one bill, the economics of running the hardware yourself and the question of how usage is counted are four other pages, each with its own angle. None is restated here. Where this page does name a technique, it is to show how the managed shape reshapes it, not to re-explain it.
azure openai pricing: what a managed cloud price actually is
Azure OpenAI Service is the model owner's models offered as a resource inside a cloud subscription, so the thing being priced is not "a model" but "a deployment of a model". Microsoft's own product page frames it as three billing modes layered on one another. Standard, also called on-demand, is the familiar pay-as-you-go shape: you are charged for input tokens, cached input tokens and output tokens at a per-million rate. Provisioned sells throughput as reserved capacity, charged at an hourly rate per unit rather than per token. Batch accepts work into a queue that returns within a defined window — Microsoft documents a 24-hour return and a discount against Global Standard pricing — which is the same latency-for-money trade a direct batch tier makes, priced against a specific deployment type rather than against a single global rate.
The consequence for the word "pricing" is that a cloud page is answering a different question than a direct card. On a direct card, one model has one input rate and one output rate, with a couple of discount rows. On a managed page, the same model is a grid: a rate per deployment type, a discounted batch column, a higher-priority processing column, and a separate hourly product for reserved capacity. This is why the taxonomy of billing mechanisms and the shape of a resold price are two different subjects; the first tells you what a provider charges for, the second tells you how a cloud vendor wraps a model owner's inference inside its own commercial objects.
azure ai gateway pricing: the deployment type is the price
On a managed platform, the deployment type is part of the price, not a setting underneath it. Azure OpenAI distinguishes at least three throughput shapes for the same model: a Global deployment that may serve the request from capacity anywhere, a Data Zone deployment that keeps processing inside a defined data boundary such as a geography, and a Regional deployment pinned to one location. The three exist because the trade between price, latency and data residency is real — a tighter processing boundary costs more, and a wider one costs less — and Microsoft prices them separately rather than exposing a single number.
Two properties follow, and both surprise readers who arrive from a direct rate card. First, the rate is not portable across deployment types: a Global reservation does not cover a Regional deployment, and the documentation lists Global, Data Zone and Regional provisioned reservations as separate purchases. Second, quota is granted per region and defines how much capacity a subscription may create at all, which means the binding constraint on a workload can be a regional allowance rather than a price. A team that reads only the per-million rate miss both: they miss that the same model has several prices, and they miss that an allowance they do not hold cannot be bought at any price. This is the layer where a managed platform behaves least like a rate card and most like a capacity market — a distinction an aggregator's pass-through spread, the story the marketplace page owns, does not have to make.
llm hosting cost: reserved capacity versus pay-per-token
The clearest shape difference between managed and direct pricing is that a managed platform will sell you hosting — reserved capacity — and a direct API usually will not. Provisioned throughput units are units of model-processing capacity; when you create a provisioned deployment you choose how many units to hold, the platform reserves and holds that capacity for the deployment, and you are billed an hourly rate per unit whether or not the deployment is handling requests. Microsoft states the rule plainly: billing is by the number of units deployed, not by tokens consumed; a deployment that exists for part of an hour is charged pro-rata; resizing adjusts the charge immediately; and a provisioned deployment cannot be paused, so billing stops only when it is deleted.
That shape flips the economic question. Under per-token billing, cost tracks usage and an idle service is nearly free. Under reserved capacity, cost tracks time, and an idle deployment is the most expensive thing you can own — a fixed hourly charge against zero output. The lever that makes reserved capacity cheaper is utilisation: reservations bought monthly or annually, and reservations scoped to a resource group, a subscription or a whole billing account, lower the effective hourly rate, but the discount is only earned by running the capacity hot. A third property is easy to miss and matters to a budget: purchasing a reservation does not guarantee capacity on the service, and units deployed beyond the reservation are billed at the standard hourly rate. So "cheaper hosting" in the reserved sense is a bet on steady utilisation, and it is a bet the direct per-token model never asks you to place.
ai infrastructure cost: what the boundary leaves in the bill
A cloud vendor's per-token rate is a fragment of the bill, because the metering boundary is drawn around the model call and the surrounding ai infrastructure is billed somewhere else on the same invoice. Nobody would call object storage, egress bandwidth, a secrets vault, or the log-analytics workspace part of "the model price", yet all of them are consumed by a production AI workload and all of them are separate meters. Azure's own pricing page is explicit that its published figures are estimates and that actual pricing varies with the agreement type, the date of purchase and the exchange rate — a disclaimer a direct rate card does not carry, and a sign that the number on the managed page is a planning input rather than an invoice line.
There is a second, quieter boundary effect. Because the managed platform meters the model call in its own units and on its own cadence — a per-million-token grid on the page, aggregated by the platform before it reaches the invoice — a rate copied out of the page and multiplied by a token count is an estimate of one component, not of the spend. The platform may also bundle capabilities into the deployment (default content filtering, for instance) that a direct integration would buy or build separately; bundling usually means the cost is present but invisible rather than absent. The honest way to reason about ai infrastructure cost on a managed platform is therefore bottom-up: the token line is one row, and the boundary around it — region, boundary tier, storage, egress, observability — is where the rest of the money lives.
openai api cost: import the model, not the rate card
The most expensive mistake in cloud pricing is assuming a model's rate travels with the model. It does not. When a team moves a workload onto a managed deployment, three of the four things that decide the bill are properties of the platform rather than of the model. The billing mode may have changed: a workload that was purely per-token may now be more economical as reserved capacity, or the platform may only offer the model you want inside a particular deployment type. The deployment type fixes a regional rate that need not match the direct card. And the agreement on top — an enterprise commitment negotiated with the cloud vendor — sets the price actually invoiced, which the published card explicitly does not fix.
For a direct comparison, then, the unit of the comparison has to change. Comparing "the rate" of a managed deployment with OpenAI's own API card compares a resold, region- and agreement-shaped price against a single global rate, and the two are not the same object: one includes a capacity product, a quota system and an enterprise contract path; the other is a metered API. The useful comparison is by workload shape — steady interactive load, latency-tolerant batch, or bursty traffic — against the modes each platform sells, and only then by number. A reader who takes one figure from each page and ranks them has compared two different products and called the result a price.
llm token cost: why the unit rate is only a fragment
Even inside the pay-as-you-go mode, the unit rate is a fragment of llm token cost on a managed platform, because the platform splits one model into several priced columns before you ever multiply. The managed grid separates deployment type (Global, Data Zone, Regional), a priority column for higher-priority processing, and a batch column priced against Global Standard rather than against the interactive rate. Input, cached input and output are separate rows in each of those columns. A context-length tier can reprice long prompts on some models, so the headline rate may apply only below a stated window. Underneath all of it sits the counting question — the meter counts in the platform's own tokeniser — which is a separate page's subject and not repeated here.
The practical reading is that a cloud vendor's token rate is a selector, not a scalar: it is the rate conditional on a deployment type, a service tier and a context window, and changing any of the three changes the number. Cache mechanics behave the same way in the managed setting as in the direct one but are consumed through the platform's columns: the discount exists, and its size and conditions are the platform's — unlike Claude's explicitly priced cache tiers, where the cache write and the re-read are separate, published rates on the card. A comparison that fixes none of these variables and then declares a winner has, once again, compared the wrong layer — which is why the shape of the price has to be the first thing pinned, and the number the second.
how much does an llm cost on a cloud: a worked shape
"how much does an llm cost" on a managed platform has no single answer, but it has a decidable shape: match the workload to the billing mode first, then price only that mode. The table below is the decision, not a price list — the numbers belong to the vendor's own page, and to your agreement.
| Workload shape | Mode a managed platform prices it as | What dominates the bill | The trap |
|---|---|---|---|
| Steady, latency-tolerant, deferrable | Batch | discounted per-token rate | the return window; not for user-facing calls |
| Steady, interactive, low volume | Standard pay-as-you-go | tokens times rate | assuming the rate is not regional |
| Steady, interactive, high utilisation | Provisioned capacity | hours times reserved units | paying for idle hours |
| Bursty or hard to forecast | Standard pay-as-you-go | tokens times rate | reserved capacity burns on the idle floor |
| Regulated data residency | Data Zone or Regional deployment | the tighter boundary's rate | a wider reservation does not cover it |
| Committed enterprise spend | Agreement pricing | the negotiated rate | the published card is an estimate only |
Read across the rows and the pattern is consistent: a managed platform sells a mode and a boundary, and the rate is a function of both. That is why the cloud vendor's number sits at a different level of the stack than a model owner's per-token rate, where DeepSeek's low per-token rates or Groq's per-token rates are quoted against one global service. Neither shape is wrong; comparing them without matching the mode is.
What our own gateway meters, read from the source
The quota-structure point is not abstract, and we can show it on our own implementation. Our gateway is not a model host — it meters the tool traffic that fills an agent's context before a model is asked to complete anything — but it faces the same design question a managed platform faces, and it answers it the same way: the cost discipline lands on a structure of independent quotas, not on a single unit price. Read from the product source on 2026-10-08:
| Plan | MCP req/min per key | MCP team ceiling (req/min) | REST write req/min (team) | Monthly token pool | Activity-log retention | Max keys per team |
|---|---|---|---|---|---|---|
| Free | 30 | 30 | 20 | 2M | 7 days | 2 |
| Pro | 300 | 600 | 120 | 20M | 30 days | 10 |
| Teams | 600 | 3000 | 300 | 100M | 90 days | 30 |
| Enterprise | 1200 | 9999 | 600 | 200M+ | 180 days | 9999 |
Three properties of that table are the ones worth copying, whatever the numbers are on the day you read this. First, the dimensions are independent: a per-key request rate and a team-wide ceiling are different limits, and a plan sets both, so a workload can be legal per key and refused per team. Second, the limits are enforced as two sliding-window buckets, one per key and one per team, so the ceiling is a real object rather than a documented number. Third, the monthly token pool is a third, slower meter, held per team as a month-stamped counter with a retention window, read against the team's cap — a utilisation budget on top of the throughput limits.
This is the same reason a cloud vendor's unit price is not the bill. In both cases the binding constraint on a workload is a quota you hold, in a place you may not be watching, and the invoice is a separate artifact from the allowance. The plan figures above are operational limits rather than a feature comparison, and they move with plan changes, so the pricing page is the authoritative table and this page is not.
How SmartGate's limits sit next to a cloud vendor's
The two are different products, and the table below is deliberately about shape rather than about which is better, because they sit at different layers.
| A cloud vendor (managed model host) | SmartGate (algorithm gateway) | |
|---|---|---|
| What is sold | Inference of a model plus reserved capacity and region | Governance of the tool traffic around a model call |
| How cost is charged | Per token, per reserved unit-hour, or by agreement | A tiered platform fee with a capped savings share |
| What limits you | Per-deployment quota, regional allowance, throughput units | Per-key and per-team request ceilings, a monthly token pool |
| What controls spend | Deployment type, mode, utilisation, agreement | Budget count and record, per-plan ceilings |
The shapes rhyme where it matters. A cloud vendor bounds a workload with several quotas — a regional allowance for capacity, a per-deployment throughput limit, a monthly agreement — and so does the gateway, with a per-key rate, a team ceiling and a token pool. Where they differ is what the money buys: the vendor sells the inference itself, so its price scales with the model; the gateway sells the layer that decides whether a call should happen and with which key, and its fee is a platform fee with a share that starts only once it has measurably saved you something, not a markup on every token. A team already committed to a cloud vendor for inference can still be exposed on the tool side, which is exactly the gap the gateway fills.
How to get started
The first move costs nothing and prevents every downstream mistake: identify the mode and the boundary before the number.
- Write the billing mode next to each candidate. Standard, provisioned, batch or agreement — one label each, and mark any candidate that is really a mix.
- Record the deployment type and region. A Global rate, a Data Zone rate and a Regional rate are three prices for one model; note which one your data-residency requirement forces.
- Decide whether hosting is on the table. If utilisation is steady and high, reserved capacity may beat per-token; if it is bursty, the idle floor usually decides it against you.
- Compare only within a mode and a boundary, then price the winning shape against the vendor's own current page — the published figures are estimates, and your agreement is the real rate.
- If your exposure is the tool layer rather than the model, start free and watch one agent call end to end, then confirm which tier your real volume needs on the pricing page.
Frequently Asked Questions
Why can't I compare an Azure OpenAI rate with OpenAI's own rate directly?
Because they are different products wearing similar clothes. Azure OpenAI is a managed deployment inside a cloud subscription: its rate depends on the deployment type and region, it adds an hourly reserved-capacity product, it is bounded by subscription quota, and Microsoft's published figures are estimates that depend on your agreement. OpenAI's own card is a single metered API. The two are comparable only after you match the billing mode and the boundary.
What is a provisioned throughput unit, in plain terms?
It is a unit of reserved model-processing capacity that you hold for a deployment. You are billed an hourly rate per unit whether or not requests arrive, the charge is pro-rated for partial hours, and a provisioned deployment cannot be paused, so billing stops only when you delete it. It is cheaper than pay-as-you-go only when you keep the reserved capacity busy.
Does a reservation guarantee that the capacity exists?
No. A reservation lowers the effective hourly rate for units within its scope, and units deployed beyond it are billed at the standard hourly rate, but purchasing a reservation does not reserve capacity on the service. Treat it as a discount mechanism, not a capacity guarantee.
Why does the same model have several prices on one cloud page?
Because the deployment type is part of the price. Global, Data Zone and Regional deployments trade price against data-processing boundaries and latency, and the platform prices them separately; batch and higher-priority processing add further columns. A single headline rate is therefore conditional on a type, a tier and a context window.
Is quota the same as a spending limit?
No, and the difference is the point. Quota governs how much throughput a subscription or deployment may use — tokens or requests per minute, sometimes granted per region — and it can bind even when you are willing to pay more. Spending tracks what you actually consume. On a managed platform you can hold quota you never bill against, and bill against capacity you never reserved.
Limitations
This page is about the shape of managed pricing, not a price list, and it deliberately copies no third-party rate it could not verify on the day it was written. Microsoft's figures are, by the platform's own statement, estimates whose actual value depends on the agreement, the purchase date and the exchange rate, so any specific number belongs on the vendor's own page and in your contract, not on this one. Where a rate is not shown here, that is a decision not to reproduce a number that moves, not a claim that the number does not exist.
The angle is also bounded on purpose. This page does not explain how to read the columns of a direct rate card, does not work through the arithmetic of a single bill, does not price self-hosted hardware, and does not describe the take rate of an aggregator or marketplace — each of those is a sibling page in this cluster, linked once above, and none is restated here. A reader who needs a decision should take the workload shape to the vendor's own current page and re-check it on the date they plan to spend against it. Finally, the plan limits quoted on this page are our own product facts, read from the product source and re-verified on the date above; they change with plan changes, so the pricing page remains authoritative and any budget should be built from it rather than from a copy.
Sources
- Microsoft's provisioned throughput billing documentation, for hourly per-unit charging, the no-pause rule, pro-rated partial hours, immediate resize effects, and the point that a reservation is a discount rather than a capacity guarantee — Microsoft Learn, "Provisioned throughput unit (PTU) costs and billing": https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/provisioned-throughput-onboarding (checked 2026-10-08).
- Microsoft's Azure OpenAI Service pricing page, for the three billing modes (Standard on-demand, Provisioned, Batch), the batch discount against Global Standard pricing and its return window, the Global / Data Zone / Regional deployment types, and the statement that published prices are estimates depending on agreement, date and exchange rate: https://azure.microsoft.com/en-us/pricing/details/azure-openai/ (checked 2026-10-08).
- Microsoft's quota and limits documentation, for quota being granted per region and bounding the throughput a subscription or deployment may create: https://learn.microsoft.com/en-us/azure/ai-foundry/openai/quotas-limits (checked 2026-10-08).
- The gateway section is our own implementation, read read-only on 2026-10-08 from the product
repository at the revision recorded in this project's
pipeline_results.json: the plan rate-limit catalog inbackend/smartgate/core/plan_entitlements.py, the two-bucket sliding-window limiter inbackend/smartgate/core/rate_limiter.py, and the month-stamped token bucket inbackend/smartgate/modules/budget_guard/__init__.py, with the plan catalog values read fromconfig/plan-catalog.ts.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice
matcher pinned 0 of the 7 planned sections for this page: rule A found no unique symbol in the scanned
repository for any section keyword, and the remote candidate fallback returned only score-ranked
near-misses, which a page must not dress up as a pin. Every section above is therefore written from
public sources — the platform's own documentation and pricing page, cited with a checked date — which
is the house rule for an unpinned section: sourced, never invented. The one exception is "what our own
gateway meters, read from the source": our own plan rate-limit catalog, sliding-window limiter and
monthly token bucket, read read-only on 2026-10-08 from
backend/smartgate/core/plan_entitlements.py, backend/smartgate/core/rate_limiter.py and
backend/smartgate/modules/budget_guard/__init__.py, stating the per-plan request ceilings, the two
independent throughput dimensions, and the month-stamped token pool. The section keyword quoted above
each heading is this project's own measured pool phrase, not a code symbol, and every one of the seven
carries a measured search volume above zero. No code, batch fingerprints, auction data or internal
hosts appear in the text.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|
0 of 7 sections pinned; 7 of 7 sections came back no-slice, 0 abstentions.