LLM Pricing: A Taxonomy of How Providers Charge
LLM pricing is not one rate; it is a small family of billing mechanisms that providers mix in different proportions. The four that cover almost every current API are per-token billing with an asymmetric input and output rate, a discounted rate for cached input, a time-based discount for batch work, and a context-length tier that reprices long prompts.
Short answer: LLM pricing is not one rate; it is a small family of billing mechanisms that providers mix in different proportions. The four that cover almost every current API are per-token billing with an asymmetric input and output rate, a discounted rate for cached input, a time-based discount for batch work, and a context-length tier that reprices long prompts. Gateways and marketplaces then add a markup, a spread or a subscription seat on top. So the first question about any price is not "how much" but "which mechanism am I actually buying".
Key takeaways
- A price is a mechanism, not a number. The same headline rate means per-token billing on one page and a reseller's blended rate on another, and the two behave differently under load.
- Input and output are priced asymmetrically on purpose. Generated tokens cost more because decoding is sequential while reading a prompt is parallel; the ratio, not the absolute rate, is what a workload feels.
- Time is a pricing input. Batch and asynchronous tiers sell the same tokens at a discount because they let the provider schedule the work on its own hardware.
- Length is a pricing input. Context tiers reprice a request above a stated threshold, so "the rate" on a card is often only the under-threshold rate.
- Every layer between you and the model can take a cut. Gateway markups, marketplace take rates and subscription seats sit on top of the provider's own rate.
- Do this first: write down which mechanism each candidate charges by before you compare a single figure, then compare only candidates that bill the same way.
Most confusion about llm pricing comes from comparing an implementation detail with a business model. A token rate and a monthly seat are both called "the price", and neither is wrong; they simply answer different questions. A page that leads with one number invites a reader to copy it into a spreadsheet, where it is silently reinterpreted as a different mechanism and the comparison quietly stops being like-for-like.
The vocabulary is a useful clue. A vendor that quotes a rate per million tokens is metering usage; one that quotes a flat monthly fee per seat is not metering at all; one that quotes a price "from" some floor is quoting a tier boundary, not a rate. Reading those signals is most of the work, and the arithmetic only starts after the mechanism has been identified — which is why the cost arithmetic itself is kept on a sibling page and not repeated here.
This page is the centre of a cluster about what AI costs, and it maps the mechanisms rather than the arithmetic. Reading the columns of one card, turning a card and a request into a bill, counting the usage those rates apply to, and costing the hardware that runs a model yourself are four separate skills; each has its own page in this cluster, linked where the section below actually needs it, and none of them is restated here.
llm pricing: the mechanisms behind one label
At least six mechanisms travel under the heading of llm pricing, and most commercial APIs mix two or three of them at once. Naming them is the fastest way to make a comparison honest:
- Per-token, split by direction. The provider counts the tokens it read and the tokens it wrote and applies two different rates. This is the dominant mechanism for hosted APIs.
- Cached input. A prefix the provider has already processed and can re-read is billed at a fraction of the input rate. It rewards stable, repeated context.
- Batch or asynchronous tiers. The same request, submitted to a queue, is billed below the interactive rate because the provider may schedule it later.
- Context-length tiers. Above a stated prompt length, one or more token classes are repriced. The headline rate then applies only inside the smaller window.
- Per-request or subscription. A flat charge per call, or a monthly seat with an allowance, where usage does not move the line item at all.
- Gateway and marketplace layers. A service that routes to a provider and bills you for the routing — a per-token markup, a per-request fee, a subscription, or a spread it keeps.
A single price list can carry three of these at once: a per-token base, a discounted batch row and a context threshold, with caching defined somewhere in a footnote. That is why "cheaper" is not a property of a vendor but of a vendor crossed with a workload. The mechanisms themselves are the subject of this page; where a card states them, column by column, is the sibling page on reading one rate card, and the multiplication that follows is the sibling on the arithmetic of a single bill.
token pricing: input and output are billed asymmetrically
Token pricing is the mechanism most people mean by "AI pricing", and its defining feature is that the two directions are priced differently. Reading a prompt is a parallel operation: every token in the input is processed at once by the same forward pass, so the provider's marginal cost per input token is low. Generating is sequential: the model emits one token, appends it, and runs again to produce the next, so the marginal cost per output token is substantially higher. A card that lists one rate for input and one for output is encoding that asymmetry, and the ratio between the two rates — not their absolute size — is what a workload actually feels. A classification job that reads a long document and writes one label is dominated by the cheap direction; a long-form generator that reads a short prompt and writes pages of text is dominated by the expensive one.
The cached-input row is the same mechanism extended in time. A prefix that a provider has computed and kept may be re-read at a lower rate on a later call, so the effective input price of a chat loop depends on how much of its context is stable. Two properties matter and are easy to miss. First, the discount is conditional: it applies on a cache hit, and whether a prefix hits is the provider's policy, not yours. Second, caching changes the shape of the input cost without changing the output cost at all, because a token the model has not generated yet cannot be cached. That is why the output-to-input ratio is the number to carry into any comparison and the input rate alone is a misleading headline.
There is a counting caveat sitting just under the price line: the meter counts in the provider's own tokenizer, so the same string can be a different number of tokens on two cards, and the counting question is separate from the pricing one — how a count is metered is its own page. The pricing consequence is narrow but real: a rate comparison that assumes identical token counts across providers is only approximately right, and the approximation is worse for text that tokenizes unevenly, such as code or non-Latin scripts.
cost per million tokens: the unit is a signal
Cost per million tokens is a convention, and the convention is informative. A rate quoted per million tokens says the vendor meters usage and expects the number to be large; a rate quoted per thousand says the same thing in a different unit; a price quoted per request or per seat says the vendor has chosen not to meter usage at all. Converting between the per-million and per-thousand forms is arithmetic, not a difference — multiply or divide by one thousand — and a card that uses one while a comparison uses the other is not in disagreement with it. What does not convert is a unit that is not a token: a per-image, per-second or per-character line needs the provider's own stated equivalence before it can enter a token comparison, and without that it stays out rather than in at a guessed rate.
The unit also tells you where the mechanism is hiding. A single "per million tokens" figure that does not separate input from output has a token mix baked into it — the vendor's mix, not yours — and it flattens the asymmetry described above into one number. A figure that is "per million tokens, cached input" is really two mechanisms printed together. Reading that correctly is easiest with a concrete card in front of you: Anthropic's cache tiers are an explicit, priced write with a lifetime, so the cached rate only applies if a prefix is re-read inside a stated interval (Claude's published cache tiers), while an automatic prefix discount applies whenever the provider decides it should (OpenAI's standard-tier card). Both can print a similar cached rate and behave differently under the same traffic, which is the point the unit alone will not tell you.
ai pricing comparison: compare mechanisms, not numbers
An ai pricing comparison that starts with the numbers is comparing the wrong layer first. Line up the mechanism before the figure, because a subscription and a per-token rate are not two prices for one thing; they are two different contracts. The useful first pass is a short table:
| Mechanism | What moves the number | What it rewards | Where it hides cost |
|---|---|---|---|
| Per-token (input/output) | the counted mix of prompt and output | shaping context and output length | the output-to-input ratio |
| Cached input | cache hits and their lifetime | stable, repeated prefixes | whether a prefix ever hits |
| Batch / asynchronous | the window you accept | work that can wait | latency, not money |
| Context tier | prompt length against a threshold | short, focused contexts | the threshold row itself |
| Subscription / seat | the seat, not the usage | steady, predictable load | overage and idle capacity |
| Gateway / marketplace | the layer's own cut | convenience and routing | a spread you do not see |
Two rows in that table are increasingly common beyond the pure API vendors. A provisioned or committed-throughput tier sells reserved capacity rather than metered tokens, which flips a workload from "pay for what you use" to "pay for what you reserve" (Azure OpenAI's provisioned deployments); and a cloud marketplace wraps several vendors' models in one billing surface, so the number on the invoice is the vendor's rate plus the platform's margin (Bedrock's per-model billing). The practical rule is narrow and worth stating plainly: a comparison is only a comparison when both sides charge by the same mechanism. When they do not, the honest output is a pair of answers for a named workload shape, not a single ranking.
The blended-rate trap survives every layer. A vendor or an aggregator that publishes one figure per million tokens for a model has chosen a mix of input, cached input and output for you, and a comparison that does not let you set those shares is a comparison of the aggregator's workload. That is why the deliverable of a good comparison is a small table — mechanism, rates, tiers, and a footnote naming any mismatch — rather than a winner's name.
cheapest llm api: cheapness depends on the mechanism
Cheapest llm api is a question with a hidden variable, and the variable is the layer being priced. A provider's own rate is the bottom of the stack; everything above it is someone's margin. Three forms of that margin are common. A gateway routes a request to a provider and charges for the routing — a per-token markup, a per-request fee, or a subscription. A marketplace pools many providers behind one key and keeps a spread; its headline rate is the provider's rate plus that spread, which can be invisible when the platform resells at a blended figure (OpenRouter's pass-through spread). An aggregator competes on convenience and may resell below list on some routes and above it on others, so its ranking is a property of its routing policy rather than of the models it lists.
Underneath those layers, the provider's own mechanism still decides which workload is cheap. A per-token vendor with a low output rate is cheap for long-form generation (DeepSeek's low per-token rates); a vendor on specialised inference hardware can push the per-token rate down on the models it runs (Groq's per-token rates); and a genuinely free tier exists for low volume, with the usual caveat that its limits, not its price, become the constraint (free-tier LLM APIs). None of those comparisons is meaningful until the mechanisms match, which is why "cheapest" is a question about your load shape as much as about a vendor. The same reasoning applies in reverse when a flat fee looks expensive: a subscription can be the cheapest option precisely because it stops metering, once the workload is steady enough to consume the allowance.
token cost calculator: what it can and cannot model
A token cost calculator models exactly one mechanism, and modelling it well is more useful than pretending to model more. For per-token billing the arithmetic is two multiplications: the input count times the input rate plus the output count times the output rate, with a third term if the service reports cached input separately. A calculator that does this honestly needs three inputs — the counts, the rates, and the unit — and a stated date, because rates move. Given those, its output is a defensible estimate of one request or one batch.
What a calculator of that shape cannot model is everything outside the per-token mechanism. It cannot price a subscription seat without assuming a volume and calling the result a rate; it cannot price a batch tier for a workload that must answer interactively; it cannot see a gateway's markup or a marketplace's spread, because those live above the rates it was given; and it cannot apply a context tier unless it has the threshold and knows which token classes the tier reprices. It also cannot model the token traffic that surrounds the call — retries, tool calls that refill the context, a failed request the provider still bills — which is the difference between the cost of one completion and the cost of running an agent. The one adjacent cost it clearly cannot reach is the price of the hardware itself, which belongs to the compute side of the question rather than to a rate card. A calculator is a good instrument for the mechanism it models and a misleading one for any other, which makes "which mechanism" the first input it should ask for.
llm cost calculator: the four values our own table multiplies
A cost calculator is only as honest as the table behind it, so it is worth showing one in full — our
own. The table is a local JSON file, read here from origin/main on 2026-10-08, and it holds four
values per model: input cost per token, output cost per token, maximum input tokens and maximum
output tokens. The unit is one token, never "per 1K", so there is no hidden division between the
stored rate and the arithmetic that uses it. The five rows at that revision read:
| Model | Input (USD / token) | Output (USD / token) | Max input | Max output |
|---|---|---|---|---|
gpt-4o |
2.5e-6 | 1.0e-5 | 128,000 | 4,096 |
gpt-4o-mini |
1.5e-7 | 6.0e-7 | 128,000 | 16,384 |
claude-3-5-sonnet-20241022 |
3.0e-6 | 1.5e-5 | 200,000 | 8,192 |
claude-3-haiku-20240307 |
2.5e-7 | 1.25e-6 | 200,000 | 4,096 |
deepseek-chat |
1.4e-7 | 2.8e-7 | 65,536 | 8,192 |
Three properties of that table are the ones a reader should copy, whatever the numbers happen to be on the day they read this. First, the per-token unit is deliberate: it removes an implicit thousand from every calculation, and it makes the moment a rate is wrong a wrong number rather than a wrong scale. Second, the tokenizer is chosen per model family — the counter uses one encoding for the GPT-4o and DeepSeek families and another otherwise — which is the mechanism behind the counting caveat above; a single character count would not be comparable across those rows. Third, and most importantly for a taxonomy, the table is a snapshot, not a quote. It is a local, diffable file that can be edited in a commit, which means a rate change is a reviewable diff rather than a deploy, and it equally means the table is only as current as its last edit. Our own rule applies to ourselves: these figures carry the date they were read, and any budget built on them should re-read the file rather than trust this page.
How SmartGate's platform pricing sits next to these mechanisms
SmartGate does not sell tokens, so it does not fit the per-token row at all. It prices a platform over the tool traffic that fills an agent's context before a model is ever asked to complete it, and it publishes operational limits rather than rates:
| Plan | Monthly token cap | MCP requests/min per key | Audit-log retention | Max keys per team |
|---|---|---|---|---|
| Free | 2M | 120 | 7 days | 2 |
| Pro | 20M | 300 | 30 days | 10 |
| Teams | 100M | 600 | 90 days | 30 |
| Enterprise | 200M+ | 1200 | 180 days | 9999 |
Read against the mechanisms above, those columns are not rates — they are ceilings and retention windows that decide what a workload is allowed to do. The billing model is the part worth stating plainly, because it is where a platform can avoid behaving like a markup: pay for the platform, and share only when it saves you something — the share begins after a measured saving threshold and the Pro total is capped, so the fee does not grow with every call the way a per-token markup would. The caps, per-key rates, retention windows and key counts change by tier, so the pricing page is the authoritative table and this page is not. Where the same mechanisms are read as a spend report rather than as a rate card, the question moves from pricing to budgeting, and the ceilings here are only half of that picture.
How to get started
The first move costs nothing and is the one that prevents every downstream mistake: identify the mechanism before the number.
- Write the mechanism next to each candidate. Per-token, cached-input, batch, context-tier, subscription, or gateway — pick one label per vendor, and mark a vendor that mixes several.
- Compare only within a mechanism. Two per-token vendors can be ranked at a stated mix; a per-token vendor and a subscription cannot be ranked without a volume assumption you should state explicitly.
- Convert every unit to one form. Move per-thousand to per-million by multiplying by one thousand, and leave per-image or per-request rows out until the provider states an equivalence.
- Set your own mix before reading any blended figure. Input, cached input and output in the proportions your workload actually produces is the only mix that decides the ranking.
- If you route tool calls through a gateway, start free and watch one call end to end, then confirm which tier your real volume needs on the pricing page.
Frequently Asked Questions
Is LLM pricing the same as token pricing?
No, and that is the most common mix-up. Token pricing is one mechanism: the provider counts input and output tokens and bills two rates. LLM pricing is the whole family, including cached input, batch tiers, context-length tiers, flat subscriptions and the gateway markups layered on top. A vendor can price tokens and still not be compared fairly with another vendor that sells seats.
Why is output more expensive than input?
Because generating a token is sequential and reading a prompt is parallel. Reading processes the whole prompt in one forward pass, so the marginal cost per input token is low; generating emits one token, appends it and runs again, so marginal cost per output token is higher. That asymmetry is structural, which is why the output-to-input ratio matters more in a comparison than either rate alone.
Do batch discounts mean the tokens are different?
No, the tokens are the same; the timing is not. A batch or asynchronous tier accepts work into a queue so the provider can schedule it on hardware it would otherwise idle, and passes part of that saving on as a lower rate. The mechanism is a trade of latency for money, so it only helps a workload that can genuinely wait.
Where do gateway and marketplace fees fit?
Above the provider's rate. A gateway charges for routing, usually as a per-token markup, a per-request fee or a subscription; a marketplace keeps a spread on the providers it pools. Those layers are real costs that a per-token calculator cannot see unless someone hands it the markup, which is why the cheapest-looking listing is not automatically the cheapest stack.
Can a single number tell me what a provider charges?
No. One figure per million tokens has a token mix baked into it, and a figure that does not separate input, cached input and output hides the asymmetry the whole mechanism is built on. Treat any single blended rate as a summary of someone else's workload, and set your own mix before you compare.
Limitations
This page is a taxonomy, not a price list, and it deliberately copies no third-party figure it could not date. The mechanisms are described generically; real vendors mix them in ways that a six-row list only approximates, and a provider's own card always wins over any summary, including this one. Where a card is not shown here, that is a decision not to reproduce a number that moves, not a claim that the number does not exist.
The taxonomy also stops at the boundary of money. It does not compute the cost of a single request, does not explain how usage is counted and attributed, does not price self-hosted hardware, and offers no savings checklist or quota design — those are separate pages in this cluster, linked through the sections above, and each owns its own angle. A reader who needs an actual figure should take the mechanism they identified here to the provider's own table, and re-check it on the date they plan to spend against it.
Finally, the plan limits quoted on this page are product facts read from the product source and re-checked against the public pricing page on the date above; they move with plan changes, so the pricing page remains authoritative and any budget should be built from it rather than from a copy.
Sources
- OpenAI API pricing and Batch documentation, for the automatic prefix-caching discount and the asynchronous tier: https://openai.com/api/pricing/ and https://platform.openai.com/docs/guides/batch
- Anthropic Claude API pricing and prompt-caching documentation, for the explicit cache-write multipliers and the cache lifetime: https://docs.anthropic.com/en/about-claude/pricing and https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching
- Google Gemini API pricing, for the per-unit conventions and context tiers a card may carry: https://ai.google.dev/gemini-api/docs/pricing
- Product behaviour and the plan table: read read-only from the product source at the revision
recorded in this project's
pipeline_results.json, with the plan figures re-verified against the live /pricing page on 2026-10-08. - The calculator section is our own implementation, read on 2026-10-08 from
backend/smartgate/modules/budget_guard/model_prices.jsonandmodules/budget_guard/algorithm.py(origin/main): the four values held per model with the rates quoted in that section, the per-token unit, and the per-family tokenizer choice.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice
matcher pinned 0 of the 7 planned sections for this page: rule A found no unique symbol in the
scanned repository for any section keyword, and the remote candidate fallback returned only
score-ranked near-misses, which a page must not dress up as a pin. Every section above is therefore
written from public sources — the providers' own cards, cited with a checked date — because the house
rule for an unpinned section is sourced, never invented. The one exception is "llm cost calculator:
the four values our own table multiplies": our own price table and counter, read read-only on
2026-10-08 from backend/smartgate/modules/budget_guard/model_prices.json with
modules/budget_guard/algorithm.py, stating the four values per model, the per-token unit and the
per-family tokenizer choice. The section keyword quoted above each heading is this project's own
measured pool phrase, not a code symbol, and every one of the seven carries a measured search volume
above zero. No code, batch fingerprints, auction data or internal hosts appear in the text.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|