Inference Cost: GPU Hours, Batching and the API Crossover
inference cost is the cost of the compute that produced the tokens, so it is GPU hours divided by the tokens those hours actually serve — not the per-token price printed on an API card. Utilization, batching and the prefill/decode split move that quotient long before any model choice does. This page works the compute-side account and finds where self-hosting and a managed endpoint cross.
Short answer: inference cost is the cost of the compute that produced the tokens, so it is GPU hours divided by the tokens those hours actually serve — not the per-token price printed on an API card. Utilization, batching and the prefill/decode split move that quotient long before any model choice does. This page works the compute-side account and finds where self-hosting and a managed endpoint cross.
Key takeaways
- A rate card is a retail price; an inference bill is a cost of goods. The provider's per-token figure already folds in the utilization it pools across tenants, so it is not what your own GPU delivers.
- A rented accelerator bills for every hour it exists. Inference cost is a fixed cost divided by a variable output, and the divisor decides which side wins.
- Utilization is the whole game. Measured fleet utilization in production sits near a quarter of full load, which multiplies the effective cost per token by roughly four against a busy card.
- Batching is the dominant lever. Iteration-level scheduling delivered 36.9x throughput at equal latency on a 175B model in Orca, and up to 24x against a naive pipeline in vLLM; static batching leaves the silicon at 30-40 percent.
- Prefill and decode are opposite workloads — compute-bound and memory-bandwidth-bound — and the serving schedule is what reconciles them.
- Do this first: measure your own tokens per hour at the concurrency you actually run, then price the accelerator against that number — everything else is arithmetic on top of it.
inference cost: the bill is a compute bill
An API card quotes a price per token. An inference bill is something different: the cost of the compute that produced those tokens, built from two quantities — the cost of an accelerator-hour, and the number of tokens that hour actually served. Write it as one line and the field becomes readable:
cost per million tokens = (cost per accelerator-hour x 1,000,000) / (tokens served per hour).
The numerator is quotable. The denominator is not: it depends on the model, the batch, the context length, the quantization and, above all, on how much of the hour the card was working. That is why "inference cost" and "API price" are different objects. The API price is a vendor's selling number, already averaged over every tenant it serves. The inference cost is your cost of goods — and the moment you run the silicon yourself, no one is pooling your idle hours with anyone else's.
Two facts frame the rest of the page. A rented accelerator bills by the hour whether or not a token crosses it, so the meter is the clock; and tokens per hour is a property of the serving configuration and the traffic that fills it, not of the model. A rate card collapses both into one number.
| Quantity | Where it lives | Who can look it up |
|---|---|---|
| Cost per accelerator-hour | the GPU cloud's price list | anyone, published |
| Tokens served per hour | your serving stack's metric | only you, and only if you measure it |
| Utilization | tokens served divided by the card's capacity | only you, over a window long enough to include the quiet hours |
| Price per token | the model provider's card | anyone, published — but it is their quotient, not yours |
The last row is the source of most bad comparisons. A provider's per-token price is its own cost per accelerator-hour divided by the tokens per hour it achieves at pooled utilization, plus a margin; the accounting behind that price — which tokens are counted, at which rate — is the subject of the billing arithmetic behind a token-priced API. This page stays a layer below it, on the compute side, where the quotient is set.
ai inference cost: accelerator hours and utilization
Start from the fixed cost, because it is the part teams under-count. A single H100 on a specialist cloud at a representative $2.50 per GPU-hour costs about $1,825 in a 730-hour month: 2.50 x 730. That bill is identical whether the card served four billion tokens or forty million, and the gap is the whole argument. A per-token endpoint charges nothing while traffic is quiet; a rented GPU charges for the quiet hours too.
Utilization is the ratio of what you served to what the card could have served. A quantized 70B-class model on one H100 saturates near 3.9 billion output tokens a month at full tilt, so divide the monthly bill by that:
- at full load, $1,825 / 3.9 billion tokens = about $0.47 per million tokens;
- at the 23 percent average utilization reported across real deployments, the same card serves about 0.9 billion tokens and the effective cost is about $2.03 per million tokens;
- at 10 percent load it is roughly $4.68 per million tokens.
The card did not change between those rows; only the divisor did — which is why an idle GPU is the fastest way to lose the cost argument. Real traffic is not flat, and that fixes the divisor in practice: a service sized for its peak hour cannot exceed the ratio of average to peak load in average utilization, and a user-facing endpoint holds headroom above peak for latency, so the honest figure is lower still. The exception is a batch job, which can approach full load for the hours it is paid for — which is why offline pipelines are where self-hosting most often pays.
The counterweight is that the divisor is also what a provider sells you: pooling thousands of tenants onto shared cards is a utilization service, absorbing your idle hours by filling them with someone else's work. That structural advantage is why the comparison must run against your own measured tokens per hour, not the card's peak. What a usage meter records per call is the input to that measurement: without a per-call count you have no divisor, only an invoice.
inference pricing: batching and the prefill/decode split
If utilization sets the divisor, the serving schedule is what raises it. The single largest lever on the cost of a token is how many sequences the card works on at once, and that is a scheduling decision, not a model one.
Two regimes. Static batching runs a group of requests together to completion, so the batch finishes only when its longest member does, short requests hold idle slots and new work waits behind the group — leaving the accelerator at roughly 30 to 40 percent on short jobs. Iteration-level (continuous) batching reschedules at every decode step: a finished sequence leaves and a waiting one takes its slot. Orca introduced iteration-level scheduling in 2022 and reported 36.9x throughput at equal latency on a 175-billion-parameter model; vLLM paired it with a paged key-value cache and reported up to 24x a naive pipeline, cutting cache fragmentation from 60-80 percent to under 4. Those numbers decide whether a card serves one token at a time or hundreds.
Batching works because the two phases of generation have opposite hardware profiles:
- Prefill processes the whole prompt in one pass. Every token is known up front, so it is a large parallel matrix multiply — compute-bound, efficient when the tensor cores are fed.
- Decode produces one token per sequence per step. The arithmetic is tiny, but the entire weight matrix and every sequence's key-value cache stream from high-bandwidth memory each step, so decode sits in the memory-bound region of the roofline. Batching many decode streams into one step amortizes the weight read across all of them, which is why throughput rises with the batch.
The catch is that the batch is two knobs welded together: turn it up and the card serves more tokens per second while every token slows down, because each decode step streams more cache. That throughput-versus-latency tradeoff is a cost decision as much as a latency one, and the techniques that manage it protect one phase from the other. Chunked prefill slices a long prompt into pieces interleaved with decode steps so a big prefill does not stall every open stream. Prefill/decode disaggregation runs the two phases on separate pools tuned for their own bottleneck; DistServe reported 7.4x more requests, or 12.6x tighter latency targets, at the same goodput. Both raise the serving stack's engineering cost, itself a line in the account. None of this changes the prompt; it changes how many tokens an hour the same card can deliver — the denominator of the cost of a token.
cost per million tokens: the quotient and what it hides
The per-million figure is the unit everyone quotes, so it is worth being explicit that it is a quotient rather than a property of a model. Multiply the hourly rate by a million and divide by the tokens the hour delivered:
$ per million = rate per accelerator-hour x 1,000,000 / tokens per hour.
Run it on the same H100 at $2.50 an hour. At full tilt that card delivers about 5.3 million tokens an hour for a quantized 70B-class model, so the unit cost is $2.50 / 5.3 = about $0.47 per million. Drop the batch — a latency-sensitive chat service that refuses to queue, a long-context workload whose attention step is the bottleneck — and delivered throughput can halve; the unit cost doubles with it. Nothing about the model moved.
Three inputs move the quotient, in order of impact:
- Utilization. It scales the denominator linearly, and it is the input teams control least and measure most rarely.
- The serving configuration. Continuous batching, chunked prefill, prefix caching and a paged cache together decide how many sequences are resident, which is the throughput multiplier.
- Precision and model size. FP8 quantization roughly doubles throughput at minimal quality loss, and INT4 can roughly triple it while shrinking the weights enough to fit a smaller, cheaper card. On a large model that can halve the unit cost or better.
Two honesty notes belong with the formula. It prices only the accelerator: weight storage, the KV-cache spill path, egress and ops time are separate lines that routinely add a multiple of the raw hardware cost. And it divides by the traffic you actually receive — a card sized for a peak it rarely sees has a real unit cost of a peak-sized bill over an average-sized workload. The published per-token cards a buyer compares against — the published per-token price cards — are each a provider's version of this quotient, computed at a utilization the buyer never sees.
ai token pricing: why one model carries several prices
A single model usually ships with more than one price per token, and the differences are not arbitrary. They are signals about the compute each tier consumes, read from the supply side.
Standard, on-demand pricing is the pooled retail rate. It assumes the provider keeps its cards busy by sharing them across tenants, so the price can be low without losing money on your quiet periods — you are renting a slice of an already-saturated card.
Batch pricing, typically half the standard rate, is priced around filling capacity the provider has already paid for. Asynchronous work that can wait lets the provider pack the schedule; the discount is a share of the utilization it recovers.
Provisioned or committed throughput is priced like the accelerator-hour it reserves. Here the provider stops pooling and holds capacity for you, so the price stops looking like a token rate and starts looking like rent — a fixed cost for a floor of tokens per minute.
Cached-input pricing is the cheapest tier and the most literal: a repeated prefix is not reprocessed, so the provider does not pay to recompute it, and it charges roughly a tenth of the input rate to pass that saved compute on. Caching is a compute optimization wearing a price tag.
Read that way, a price list is a menu of utilization deals rather than a set of unrelated numbers, each tier a different assumption about how full the card will be. That is also why the tiers are not interchangeable for a buyer: a workload that is constant and predictable can be priced like reserved capacity, one that is spiky belongs on a pooled rate, and one that repeats the same long prefix earns the lowest tier. Matching the workload to the tier means knowing which tokens each call sends — attributing spend to one agent run is where that starts.
gpu cost: what an hour of accelerator actually rents for
The numerator of everything above is the hourly rate, so read the published market rather than a single headline. The spread between providers for the same silicon is larger than the spread between adjacent models, and the cheapest visible number is rarely the cheapest usable system.
| Instance (one GPU) | On-demand per GPU-hour | Source and shape |
|---|---|---|
| H100 SXM 80GB | $3.29 | Lambda list, 1x self-serve config |
| H100 PCIe 80GB | $2.49 | Lambda list, 1x config |
| A100 SXM 80GB | $1.79 | Lambda list, 8x config, per GPU |
| A100 SXM 40GB | $1.29 | Lambda list, 1x config |
| GH200 96GB | $1.49 | Lambda list, 1x config |
| A6000 48GB | $0.80 | Lambda list, 1x config |
| A10 24GB | $0.75 | Lambda list, 1x config |
Those are Lambda's own published rates, checked 2026-10-03; multi-GPU configurations list lower per-GPU rates because the instance is priced as a whole. The wider market spans roughly $1.50 to $3.00 an hour for an H100 on specialist clouds, with hyperscaler reserved capacity materially higher: AWS prices a p5.48xlarge capacity block — eight H100s — at $5.19 per accelerator-hour in US East (N. Virginia), read 2026-10-03. A dedicated 24/7 H100 therefore lands between about $1,000 and $5,000 a month depending on the provider tier, before a second card for failover.
The rate is the beginning of the cost, not the end. The card also bills for hours it did not work, so the effective rate is the list rate divided by utilization; a production service usually holds a second instance, roughly doubling the bill before it serves an extra token; and the serving stack — inference server, driver and CUDA management, autoscaling, monitoring, a regression test per model update — is engineering time that routinely adds a multiple of the raw hardware cost. Weight storage, the KV-cache spill path and egress are further charges on some providers and hidden on others, and power, cooling and depreciation belong to the same list for owned hardware. The cheap number in a price comparison is the rate; the number that decides the argument is the fully loaded cost of an hour that actually served traffic — a better reduction target, which is what reducing the tokens an application sends goes after on the workload side.
gpu pricing: where self-hosting crosses a managed API
Now the two sides sit on one axis. Self-hosting is a fixed cost; an API is a variable one. The crossover is the volume at which the fixed bill equals what the API would have charged:
break-even tokens per month = monthly GPU bill / blended API price per token.
With the H100 at $1,825 a month and a representative open-model API at $0.90 per million tokens, that is about 2.0 billion tokens a month — and only if the card is actually busy serving them. Add the second card for failover and the calculation roughly doubles to 4 billion. Below the line the API wins, and the API's zero-idle pricing is doing work the fixed bill cannot.
The decisive variable is the utilization a renting team can sustain, and the published evidence is unforgiving. Against mid-priced providers a rented accelerator breaks even when it stays busy for roughly a fifth to most of every hour, depending on the model and the cloud; against the cheapest per-token tier for the same open model, a single rented card typically does not break even at any volume it can physically serve, because the discount provider runs at a utilization no single tenant matches. Two regularities are worth keeping:
- Cheap GPU clouds roughly halve the utilization a self-hosted setup needs, compared with hyperscaler on-demand rates — a bigger effect than the choice of model.
- Batch workloads clear the bar that interactive ones miss. A nightly job holds a card near full load for the hours it pays for; a service sized for peak traffic cannot.
Against a frontier model the crossover arrives far sooner, in the low millions of tokens a day, simply because the alternative is expensive — a quality trade as much as a hosting one, since the honest comparison is the open model you would self-host against the tier of API you would otherwise buy.
The pattern that survives real traffic is hybrid routing by economics rather than conviction: steady high volume on owned or reserved capacity where the batches fill, spiky or frontier-quality work on a pooled API, and one instrumented record of which call went where. Deciding which calls may spend on which path is the control half of the same problem, and capping a team's monthly token spend is where the ceiling is enforced rather than merely observed.
What a cost guard actually answers, read from our own module
An inference-cost system ends in a decision — allow, or stop — and the shape of that decision is a
design choice. Ours is visible in two files, read on 2026-10-07 from
backend/smartgate/modules/budget_guard/models.py and the counter beside it.
- Three verbs, one record each:
check,count,record. Checking asks what the budget allows, counting asks what a piece of text costs, and recording writes what was actually spent. Keeping them separate is what lets a caller ask before spending rather than reconcile afterwards. - The request carries everything the guard needs to be authoritative: the model, the text (or the prompt and completion token counts directly), a cost, the team's monthly limit, an entitlement limit and an optional client-side budget. A guard that has to look up any of those itself has already lost the race with the request it is guarding.
- The answer is four numbers and a boolean: used, limit, remaining, exceeded. That is the whole contract — deliberately not a cost breakdown, because the caller's only decision is whether to proceed and how much headroom is left if it does.
- The cost of a call is priced by a table, and counted by a tokeniser (input and output priced separately, encoding chosen per model family). Both halves are replaceable configuration, so a wrong price is a data fix and a wrong count is a code fix — and you should know which one you are looking at before you argue about a bill.
Where SmartGate fits
The account above is about the compute a token costs. A large share of a real inference bill is not the completion at all — it is the tool traffic that fills the context before the completion is requested, and that traffic is governed one layer earlier. SmartGate is the gateway at that layer, and its numbers are operational limits rather than per-token rates:
| Plan | Monthly token cap | MCP requests/min per key | Audit-log retention | Max keys per team |
|---|---|---|---|---|
| Free | 2M | 120 | 7 days | 2 |
| Pro | 20M | 300 | 30 days | 10 |
| Teams | 100M | 600 | 90 days | 30 |
| Enterprise | 200M+ | 1200 | 180 days | 9999 |
The billing model is the part worth stating plainly: pay for the platform, and share only when it saves you something — the savings share starts after a measured threshold rather than on every call, so the fee does not scale with the tool traffic the way a per-token proxy would. Read against the compute account above, the two layers compose: the gateway decides which tool calls happen and caps the token traffic they generate, and the accelerator — yours or a provider's — pays for the completion that follows. The caps, per-key rates, retention and key counts change with the plan; the pricing page is the authoritative table, and any ceiling that has to survive a budget review should be read there rather than from this page — the same counters are what the finops view reports against.
How to get started
- Measure your tokens per hour before you price anything. Log output tokens served and the hours the card was resident; divide. An unmeasured denominator makes every comparison below noise.
- Compute the unit cost from your own numbers. Rate per accelerator-hour times one million, divided by measured tokens per hour. That is your cost of goods, and it is what a provider's per-token price can be compared against honestly.
- Model both sides at your utilization, not at peak. Take the monthly bill of the accelerator you would rent, divide by the blended per-token price of the API you would otherwise use, and check whether your real traffic clears the break-even volume.
- Route by economics. Start on a pooled API to measure your real traffic on real data with start free, then move the one steady, high-volume workload to reserved capacity only when its utilization clears the bar, keeping the long tail on the API. The pricing page is the source for the caps above.
Frequently Asked Questions
Is inference cost the same as the price per token on an API card?
No. The card is a vendor's selling price, already averaged over the utilization it pools across tenants. Inference cost is the cost of the compute that produced the tokens: the accelerator-hour rate divided by the tokens that hour actually served. The two converge only when your utilization matches the provider's.
Why does the same model cost different amounts per token on different setups?
Because the per-million figure is a quotient, not a property of the model. The denominator depends on the serving configuration, the precision (FP8 roughly doubles throughput, INT4 roughly triples it) and above all on how busy the card is. Same weights, different denominator, different unit cost.
What is a realistic utilization to plan for?
Lower than most plans assume. Measured production utilization sits near a quarter of full load, and a service sized for a peak three times its average cannot exceed one third before headroom for latency. Batch jobs are the exception.
Does batching lower inference cost or only latency?
It lowers cost, and it can raise latency. Continuous batching lifts the tokens per second a card delivers, which raises the denominator of the unit cost; the same larger batch lengthens each decode step, so per-token latency rises with it. The operating point just below the throughput knee keeps both acceptable.
Does self-hosting always beat an API above some volume?
No. Against the cheapest per-token tier for the same open model, a single rented card usually does not break even at any volume it can serve, because the provider's utilization is one a single tenant cannot match. Self-hosting pays against mid-priced providers once the card stays busy for roughly a fifth to most of every hour, and against frontier models at much lower volume.
Where does the per-token billing arithmetic fit relative to this?
One layer up. The cost of a token is what this page computes from compute; how counted tokens at input, cached-input and output rates become an invoice is the accounting on top of it, and the two compose rather than replace each other.
Limitations
This page works a model, and a model is only as good as the numbers fed into it. The GPU rates quoted are published list prices from the named providers, checked on 2026-10-03, and they move: reserved and spot terms, regional pricing and negotiated contracts all differ from the on-demand list, and a budget should recompute against the current table rather than this page.
The worked examples assume a clean, comparable workload. Real traffic carries input and output tokens at a ratio this page does not fix, prompts of widely varying length, and multimodal and storage line items priced on separate rows; a card's usable throughput for a long-context RAG pattern is lower than for a short chat pattern, and every self-hosted unit cost scales inversely with whatever throughput the traffic actually achieves.
The benchmark figures are other people's numbers, not measurements of any system discussed here. They establish the direction and rough magnitude of each effect — batching, the prefill/decode split, quantization — and they were produced on specific hardware and software revisions; a different stack at a different revision will land elsewhere.
Finally, the savings checklist, quota design and monitoring are separate pages in this cluster, linked above, each owning its angle.
Sources
-
Orca — A Distributed Serving System for Transformer-Based Generative Models, OSDI 2022, for iteration-level scheduling and the 36.9x result: https://www.usenix.org/conference/osdi22/presentation/yu
-
vLLM — Efficient Memory Management for Large Language Model Serving with PagedAttention, SOSP 2023, for the paged KV cache, the up-to-24x result and the fragmentation figures: https://arxiv.org/abs/2309.06180
-
DistServe — Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving, OSDI 2024, for the prefill/decode split and the 7.4x / 12.6x results: https://arxiv.org/abs/2401.09670
-
Roofline model (Williams, Waterman, Patterson), for the compute-bound versus memory-bound reading of prefill and decode: https://dl.acm.org/doi/10.1145/1498765.1498785
-
Lambda Cloud GPU pricing (per-GPU-hour on-demand rates): https://lambdalabs.com/cloud , checked 2026-10-03
-
AWS EC2 Capacity Blocks for ML pricing (p5.48xlarge accelerator-hour, US East): https://aws.amazon.com/ec2/capacityblocks/pricing/ , checked 2026-10-03
-
TeachMeIDEA — Self-Hosting an LLM: The Break-Even Point Against Token APIs, for the utilization band and the cheapest-tier result: https://teachmeidea.com/self-hosting-llm-break-even/
-
osFoundry — When Self-Hosting LLMs Is Actually Cheaper Than an API, for the specialist-cloud range and the batching and quantization levers: https://osfoundry.io/articles/when-self-hosting-llms-is-cheaper
-
SpendArk — Self-Hosting Llama vs OpenAI API: Break-Even Math, for the fixed-versus-variable shape and the per-million figures at 23 percent utilization: https://spendark.com/blog/self-hosting-llama-vs-openai-api/ , reporting the Harness 2025 average GPU-utilization figure used above.
-
Demand figures for this page's section keywords are this project's own measurement: DataForSEO Google Ads, United States, English, 12-month window, recorded in this project's search_volume.json and research_brief.md, measured 2026-10-03.
-
Product behaviour and the plan table: read read-only from the product source at the revision recorded in this project's pipeline_results.json, with the plan figures re-verified against the live /pricing page on 2026-10-03.
-
The section "What a cost guard actually answers, read from our own module" is our own implementation, read on 2026-10-07 from
backend/smartgate/modules/budget_guard/models.pywithmodules/budget_guard/algorithm.py(origin/main). It states only what those files state.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of 7 sections for this page (0 abstention(s), 7 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the seven section keywords, because this lane's vocabulary — inference, pricing, gpu, tokens — collides with generic billing, trend and estimate helper names across a codebase (several sections returned near-identical fallback candidate lists, none unique). So every section but one above is written from public sources, each cited with its own checked date.
The GPU rates and benchmark results are the named providers' and papers' published figures, not our commitments, and they move; a decision that depends on them should re-read the source. Product claims were read read-only from the product source at the revision the slice run recorded in this project's pipeline_results.json, and the plan figures were re-verified against the live pricing page on 2026-10-03. The section keyword quoted above each heading comes from this project's own paid measurement run, not from a third-party tool. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.