AI Model Pricing: How to Read a Rate Card, Not a Bill
AI model pricing is a document you read, not a single number you multiply. A rate card states separate input, cached-input and output rates, usually quoted per million tokens, and lays on tiers — cache writes, batch, long-context — that decide which rate a given token is billed at.
Short answer: AI model pricing is a document you read, not a single number you multiply. A rate card states separate input, cached-input and output rates, usually quoted per million tokens, and lays on tiers — cache writes, batch, long-context — that decide which rate a given token is billed at. Reading it correctly means converting the unit, identifying the tier, and knowing the conditions under which two cards are not comparable at all.
Key takeaways
- There is no "price of a model". There is an input rate, an output rate and a cached-input rate, and the ratio between them is what tells you which workload shape a card is built for.
- The unit is a convention, not a fact. Per-million and per-1K are the same rate, scaled by 1,000; a card that quotes one and a comparison that quotes the other are not in disagreement.
- Three tiers reprice the same token. Cache writes, batch jobs and long-context requests bill the same words at a different rate, so a card without its tier column is only half read.
- Caching is a semantics question before it is a price question. Automatic prefix caching and an explicit cache write with a lifetime are different products with the same-looking discount.
- Two cards are comparable only when the units, the cache semantics and the context tiers match. When they do not, the honest move is to say so, not to line the headline rates up.
- Do this first: take one card, copy its three rates and its tier multipliers onto a single page, and read a real workload against it before you compare it to a second provider.
ai model pricing: what a rate card actually states
A rate card is not a price for a model; it is a small matrix, and every cell answers a different question. Read it in three passes. The first pass is the three base rates: input, cached input and output. The second pass is the tier multipliers: what changes the rate without changing the text — a cache write, a batch submission, a request past a context threshold. The third pass is the unit and the scope: is the number per million tokens or per thousand, per token or per image, and does it apply to every model on the page or only to the ones in a footnote.
Most misreadings happen in the third pass. A headline that says "from $0.20" is a floor, not a rate; the model you actually call is a row further down. A card quoted per 1K is not cheaper than one quoted per million — $0.002 per 1K and $2.00 per million are the same number written twice. And a rate that is "per request" or "per generated image" is a different unit entirely, which is why an audio, image or embedding row cannot be folded into a token comparison without a conversion the provider has to state.
The first pass is where a comparison usually goes wrong. If a reader takes only the input rate, two cards can look far apart and behave far closer: a card with a low input rate and a high output rate is cheap for classification and expensive for long-form generation, and the output-to-input ratio is what reveals it. So the reading order that survives contact with a real bill is output first, then cached input, then input — the two columns you will pay most, then the column you can shrink with structure. The mechanics of turning those three rates into the cost of one specific request belong to the arithmetic of one bill; this page is about reading the card before any arithmetic starts.
token pricing: the three rates behind a single line
Token pricing is three prices sharing one card, and a line that reports only "input" is reporting the cheapest of the three. The other two are the ones with structure attached:
- Input (uncached) is the full rate on every token the model has not already processed — the prompt, the running conversation, the tool schemas, any pasted document.
- Cached input is a prefix the provider stored and re-reads on a later call. On the current cards it bills at roughly a tenth of the input rate, which makes prompt layout a pricing decision.
- Output is each token the model generates, and on almost every current card it is the expensive class — commonly five to six times the input rate.
Two properties of that trio decide how a card should be read. First, the cached-input rate is conditional: it applies only on a cache hit, and whether a prefix hits depends on how the provider manages the cache, which differs between vendors. Second, the output rate is unavoidable: you can compress an input, but you cannot cache a token you have not generated yet, so a generation-heavy workload has less room to move than a retrieval-heavy one on the same card. Reading pricing correctly therefore starts by asking which of the three classes a workload actually spends, because the same card is a different product to a chat loop and to a batch classifier.
There is also a counting caveat that is easy to miss from the price line alone: the meter counts tokens in the provider's own tokenizer, so the same string can be a different number of tokens on two cards. That is a metering fact rather than a pricing one — where a count comes from and how it is recorded is the usage-metering question — but it is why a rate comparison that assumes identical token counts is only approximately right.
model pricing comparison: how to line up two cards
A model pricing comparison is a reading exercise with one rule: line up like with like, or state that you cannot. Three columns carry almost all the information, and a fourth usually decides the answer:
| Column to compare | What it tells you | Why it decides the ranking |
|---|---|---|
| Output rate | the cost of every generated token | the largest and least avoidable line on most cards |
| Cached-input rate | the cost of re-reading a stable prefix | separates cards by how much structure can save you |
| Input rate | the cost of new context | the headline number, and often the least decisive |
| Tier column | whether a cheaper rate even applies | a batch or long-context row can reorder two cards |
Read across those columns and a comparison has three honest outcomes. It can be comparable — same unit, same cache semantics, same tier definitions — in which case the columns rank the cards for a stated workload shape. It can be conditionally comparable, where the ranking flips between a read-heavy and a write-heavy workload, so the honest answer is "depends on the mix" with the mix named. Or it can be not comparable, where a difference in billing semantics makes the headline rates describe different things. The next two sections read one card each so the columns have concrete numbers; the last section names the four conditions that make the third outcome unavoidable.
The trap that survives every comparison is the single blended number. A provider or an aggregator that publishes one "price per million" for a model has chosen a token mix for you, and that mix is almost never yours. A comparison that does not let you set the shares of uncached input, cached input and output is a comparison of the aggregator's workload, not of yours.
gpt-4o pricing: reading one provider's list card
OpenAI's published card is the canonical shape, and reading it is a matter of knowing which row you are in. The current GPT-5.6 family lists three models with the same structure and different numbers: Sol at $4.00 per million input, $0.40 cached and $20.00 output; Terra at $2.00, $0.20 and $12.00; and Luna at $0.20, $0.02 and $1.20 (OpenAI published list price, standard tier, checked 2026-10-03). Three things are worth noticing without multiplying anything. The cached-input column is exactly a tenth of the input column across all three, so on this card caching does not separate the models — the base rates do. The output column is five to six times the input column, which is the ratio a reader should carry into any comparison. And the three rows are the same product at three price points, so "which model" and "which rate" are two different questions on the same page.
The older gpt-4o row is the generational baseline most readers still anchor on, and this page does
not reproduce its exact rates: those numbers were not re-verified on the date above, so treat them as
需现场核对 — re-check the provider's own card rather than a figure copied from a comparison site.
The reading method does not change with the generation, which is the point. Whatever the row, you are
looking for the same three columns and the same footnote tiers: a Batch tier at half the standard rate
for work that can wait up to 24 hours, a long-context multiplier above a stated threshold, and a
cache-write premium on the models that support explicit caching. A card read this way answers "which
rate applies to me" for any model in the family, including one that did not exist when the card was
written.
That is also why a provider's own card is the authority and a comparison blog is a starting point. The card is the thing the invoice is computed from; everything else is a reading of it, and readings age.
claude api pricing: cache semantics change the unit
Anthropic's Claude API pricing uses the same three columns with one structural difference that changes how the card must be read: caching there is an explicit, priced operation with a lifetime, not an automatic prefix discount. A cache write bills a premium — 1.25 times the base input rate for a five-minute cache and 2 times for a one-hour cache — and a cache read then bills at 0.1 times the base input rate (Claude API pricing, checked 2026-10-03). The current family lists Sonnet 5 at $2.00 input, $0.20 cached and $10.00 output; Opus 5 at $5.00, $0.50 and $25.00; and Haiku 4.5 at $1.00, $0.10 and $5.00.
The difference is not the size of the discount; it is what the discount is attached to. On an automatic-cache card, "cached input" is a rate the provider decides when to apply. On this card it is a decision you make and pay for, with a clock: a prefix re-read inside the cache's lifetime earns the cheap rate, and one used more slowly than the lifetime pays the write premium and never reads. So a Claude card has a fourth number a reader must hold — the reuse interval of each prefix — before the cached-input column means anything. Two cards can print an identical cached rate and behave differently under the same traffic, purely because one caches implicitly and the other sells you the write.
This is the first place a comparison can honestly fail. If you line up the input and output columns of this card against an automatic-cache card, you have compared two products that happen to share a column layout. The rest of the arithmetic is identical — three rates, multiplied by counts — which is the bill-side question this page deliberately leaves to its cluster sibling. What this section adds is the reading rule: a cached-input rate is meaningless without the caching semantics printed next to it.
gemini api pricing: context tiers and the per-1K form
A gemini api pricing page is where the tier column earns its keep, because tiered-by-context pricing is the shape that most often makes a straight rate comparison wrong. The rule to look for is a threshold stated in tokens: below it, one rate; above it, a higher rate on some or all token classes, because a longer prompt costs the provider more to process. A reader who compares only the under-the- threshold rate has compared the smaller half of the card. Treat any specific Gemini row as 需现场核对 here — this page did not verify those numbers on 2026-10-03, and a threshold is exactly the kind of figure that moves between releases.
The second thing to read carefully is the unit. Google's published rate cards have historically mixed forms across their product lines — some quoted per 1,000 characters or per 1K tokens, others per million — which is not a contradiction but a convention, and it is the single most common source of a false "cheaper" verdict. Convert before you compare: a rate of $0.002 per 1K tokens is $2.00 per million; a rate of $12.00 per million is $0.012 per 1K. Once the units match, the three-column method applies unchanged. What does not convert cleanly is a per-character or per-image rate, because a character is not a token and an image is not a fixed number of them — those rows need the provider's own stated equivalence, and without it they stay out of the comparison rather than in it at a guessed rate.
So a Gemini card is read in two passes that other cards sometimes let you skip: confirm the unit, then find the context threshold and note which classes it reprices. Get those two right and the card compares cleanly; skip them and the comparison is of two numbers that were never describing the same request.
llm price comparison: when two cards are not comparable
An llm price comparison that ends in a single ranking is usually hiding one of four mismatches. Each one is a condition under which the headline rates describe different things, and each has a specific fix that is cheaper than the wrong verdict:
- Different units. Per-million against per-1K is arithmetic, not a difference — divide or multiply by 1,000 and move on. Per-token against per-request, per-image or per-second is not; those rows need the provider's stated equivalence or they do not enter the comparison.
- Different cache semantics. An automatic prefix discount and an explicit, priced cache write with a lifetime are two products. The fix is to state which one each card is and to compare the caching columns only when the semantics match.
- Different context tiers. If one card reprices past a threshold and the other does not, the rates are comparable only inside the shared range. Name the threshold and compare like-for-like inside it.
- Different tier availability. A batch rate, a committed-throughput discount or a regional price changes which rate a real workload lands on. A comparison run at the standard tier is a comparison of the most expensive way to buy each card, which may not be how anyone buys either.
When even one of those is unmatched, the honest comparison has no single winner; it has a workload shape and a set of conditions. That is not a failure of the comparison — it is the correct output of reading two documents that answer slightly different questions. The useful deliverable is a small table: two cards, three rates each, the tier multipliers, and a footnote naming the mismatch. A reader can then apply their own mix, which is the only mix that decides the ranking. Where the money actually goes once a workload is running — routing, retries, the tool traffic that fills the context before the completion — is a separate question from what the card says, and it is why a falling list price has not produced falling bills. That execution-side view is a per-execution cost breakdown, and the levers that shrink a bill before the rate is even applied are collected as token optimization techniques.
What our own price table holds, and how we count against it
Every method above ends in the same place: a number multiplied by a counted quantity. Here is our own
version of both halves, read on 2026-10-07 from
backend/smartgate/modules/budget_guard/model_prices.json and the counter beside it.
- The table is a local, diffable file with four values per model — input cost per token, output
cost per token, maximum input tokens, maximum output tokens. Read on that date:
gpt-4o2.5e-6 and 1e-5 per token (128k in, 4k out) ·gpt-4o-mini1.5e-7 and 6e-7 ·claude-3-5-sonnet-202410223e-6 and 1.5e-5 (200k / 8k) ·claude-3-haiku-202403072.5e-7 and 1.25e-6 (200k / 4k) ·deepseek-chat1.4e-7 and 2.8e-7 (64k / 8k). - The unit is one token, never "per 1K". The published cards above this section use per-1K, which is fine for a human reading a page and a trap in code: keeping the unit per token means there is no hidden division between the card and the arithmetic.
- The tokenizer is chosen per model family. The counter uses one encoding for the GPT-4o and DeepSeek families and a different one otherwise. That is the mechanism behind the warning above: two cards are not comparable by multiplying a price by a single character count, because the count itself changes with the model.
- A local table is a snapshot, so date it. Rate changes are a file diff here rather than a logic deploy — which is the property to copy — but it also means the table is only as current as its last edit. Our own rule applies to ourselves: the figures above carry the date they were read.
How SmartGate's price sits next to a rate card
A rate card prices a completion; a gateway is priced as a platform over the tool traffic that fills the context before that completion is requested. The two are different units on purpose, and reading them as if they were comparable is the same mistake as lining up per-million and per-request rows. SmartGate does not publish per-token rates; it publishes operational limits:
| Plan | Monthly token cap | MCP requests/min per key | Audit-log retention | Max keys per team |
|---|---|---|---|---|
| Free | 2M | 120 | 7 days | 2 |
| Pro | 20M | 300 | 30 days | 10 |
| Teams | 100M | 600 | 90 days | 30 |
| Enterprise | 200M+ | 1200 | 180 days | 9999 |
Read against the cards above, the columns are not rates at all — they are the ceilings and the retention windows that decide what a workload is allowed to do. The billing model is the differentiator worth stating plainly: pay for the platform, and share only when it saves you something — the share starts after a measured savings threshold and the Pro total is capped, so the fee does not scale with every tool call the way a per-token markup would. The caps, per-key rates, retention and key counts change with the plan, so the pricing page is the authoritative table, and the same rows read as a spend report rather than a rate card are the FinOps lead's view. The ceilings themselves are enforced by the gateway, and the enforcement mechanics — the entitlement read, the pre-flight counter, the refusal — are worked through in enforcing a token quota per team. The per-minute half of the same limits, and the queues behind a burst, are concurrency and rate limiting. The compute side of the same bill — what a self-hosted box costs per hour against a rented rate — is a different calculation again, kept on compute and self-hosting cost.
How to get started
- Copy one card onto one page. Three base rates and the tier multipliers, with the provider's name and the date you checked it written next to the numbers. If a row is not on the card, it does not enter the comparison.
- Convert every unit on the page to per-million tokens. Divide per-1K rates by 1,000; leave per-image and per-request rows out until you have the provider's stated equivalence.
- Mark the tiers that apply to you. Note the batch rate if your work can wait, the context threshold if your prompts are long, and whether caching is automatic or an explicit write you pay for.
- Compare at your own mix. Rank the cards for a read-heavy and a write-heavy version of your workload separately, and write down which card wins each — that pair of answers is the comparison.
Start on the free tier — 2M tokens a month, the full MCP tool surface and no card — with start free, then confirm the current plan limits against your own workload on the pricing page.
Frequently Asked Questions
Is a rate card the same as the price of a model?
No. A model has no single price. A card gives you at least three rates — input, cached input and output — plus tier multipliers that reprice the same token. Quoting one number for a model means the provider has chosen a token mix for you, which is a summary, not a price.
Why do two providers quote prices in different units?
Because the unit is a convention, not a fact. A rate per million tokens and the same rate per thousand tokens are the same number scaled by 1,000, and providers pick whichever reads well for their product. Convert before comparing; a unit difference is never a price difference.
What makes two rate cards not comparable?
Four things: a different unit, different caching semantics, a different context threshold, or a tier that only one card offers. When any of those is unmatched, the headline rates describe different products, so the honest comparison names the mismatch instead of declaring a winner.
Does a lower cached-input rate always mean cheaper caching?
No. A cached-input rate only applies on a cache hit, and whether a prefix hits depends on the caching semantics — an automatic discount applies when the provider says so, while an explicit cache write with a lifetime must be paid for and re-read inside that lifetime. The rate is a column; the semantics decide whether you ever reach it.
Should I compare cards on the input rate?
Rarely. Output is the expensive class on the current cards, at five to six times input, and it is the one you cannot cache away. A comparison that starts with output, then cached input, then input ranks cards by what a real workload actually spends.
Limitations
This page teaches a reading method; it is not a price list, and it deliberately copies no figure it
could not date. The rates quoted in the OpenAI and Claude walkthroughs are the providers' published
standard-tier list prices checked on 2026-10-03, not a commitment by us or by them, and they move:
promotions expire, regional and negotiated rates differ from the public table, and a row that was not
re-verified on that date — the gpt-4o generation and every Gemini row — is marked for on-site
checking rather than reproduced. Any comparison built from these instructions must be recomputed
against the providers' own cards before it goes into a budget.
The method also assumes a clean, text-token comparison. Cards carry rows this page does not model: image, audio, embedding and storage line items priced per unit rather than per token, committed-use discounts, and minimum-spend commitments. Each of those can move a real invoice, and none of them collapses into a per-million rate without the provider stating the equivalence. A card read perfectly is still only the card; it says nothing about how a workload is actually routed, and the difference between the two is the gap between a list price and a bill.
Finally, this page is a reading page on purpose. It does not compute the cost of one request, does not price a self-hosted model, does not explain how usage is metered, and gives no savings checklist or quota design — those are separate pages in this cluster, linked through the sections above, and each owns its own angle.
Sources
- OpenAI API pricing (standard-tier input, cached-input and output rates for the GPT-5.6 family, plus the Batch and long-context tier descriptions): https://openai.com/api/pricing/ , checked 2026-10-03.
- Anthropic Claude API pricing (base input, five-minute and one-hour cache writes, cache reads, output rates for the current family): https://docs.anthropic.com/en/about-claude/pricing , checked 2026-10-03.
- Anthropic prompt-caching documentation, for the cache-write and cache-read multipliers and the cache lifetime: https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching
- OpenAI prompt-caching and Batch documentation, for automatic prefix caching and the 50% async tier: https://platform.openai.com/docs/guides/prompt-caching and https://platform.openai.com/docs/guides/batch
- Google Gemini API pricing (per-unit conventions and the context tiers to verify on the live card): https://ai.google.dev/gemini-api/docs/pricing
- Product behaviour and the plan table: read read-only from the product source at the revision recorded
in this project's
pipeline_results.json, with the plan figures re-verified against the live /pricing page on 2026-10-03. - The price-table section is our own implementation, read on 2026-10-07 from
backend/smartgate/modules/budget_guard/model_prices.jsonandmodules/budget_guard/algorithm.py(origin/main): the four values held per model with the rates quoted in that section, the per-token unit, and the per-family tokenizer choice.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher
pinned 0 of 7 sections for this page: rule A found no unique symbol in the scanned repository for any of
the seven section keywords, and the remote symbol-candidates fallback returned only score-ranked
near-misses, which a page must not dress up as a pin. Every section above is therefore written from
public sources — the providers' own rate cards, cited with a checked date — because the house rule for
an unpinned section is sourced, never invented. The one exception is "What our own price table holds,
and how we count against it": our own implementation, read on 2026-10-07 from
backend/smartgate/modules/budget_guard/model_prices.json and modules/budget_guard/algorithm.py,
stating the four values per model, the per-token unit and the per-family tokenizer choice. The section keyword quoted above each heading is this
project's own measured pool phrase, not a code symbol, and every one of the seven carries a measured
search volume above zero. No code, batch fingerprints, auction data or internal hosts appear in the
text.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|
Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before publication. 0 of 7 sections pinned, 0 abstentions, 7 no-slice sections (all seven written from the providers' published cards and cited above).