SmartGateSmartGate

Claude Pricing: Seats, API Tokens, and the Cache Asymmetry

Claude pricing is two prices for one model family. A monthly subscription seat — Pro, Max, or a Team seat — sells an allowance of usage for a flat fee, while an API key sells the same model per token, billed separately for input, output, cache writes and cache reads.

Short answer: Claude pricing is two prices for one model family. A monthly subscription seat — Pro, Max, or a Team seat — sells an allowance of usage for a flat fee, while an API key sells the same model per token, billed separately for input, output, cache writes and cache reads. The figures that decide most real workloads are not the headline rate but two asymmetries: output is priced at about five times input, and a cache write is billed above a fresh input token while a cache read is billed at about a tenth of one.

Key takeaways

  • The same Sonnet-class model has two prices: roughly three dollars per million input tokens and fifteen per million output tokens on the API, or an allowance inside a twenty-to-two-hundred dollar monthly seat.
  • Output is priced at about five times input across the current lineup, so a workload's cost follows the length of what the model writes at least as much as the length of what it reads.
  • Caching is deliberately asymmetric: a five-minute cache write is billed above the base input rate, a one-hour write at twice it, and a cache read at roughly a tenth of the input rate.
  • A cache write only pays for itself if the prefix is read back — one reuse clears a five-minute write, while a one-hour write wants the prefix reused at least twice before it is ahead.
  • A subscription hides the per-token columns; an API key exposes every one of them, including the two cache columns a seat never itemises.
  • Do this first: decide whether your load is steady enough to sit inside a seat's allowance, and if it is not, meter the key and make the cache read rate — not the headline input rate — the number you optimise.

Most arguments about Claude pricing are really arguments about a door. The same model is sold behind two of them, and the two doors bill in different currencies: one charges a monthly fee for an allowance of usage, the other charges a rate for each token that crosses the wire. Neither is the "real" price, and a comparison that mixes them without saying which door it walked through is quietly comparing a subscription to a meter.

This page is about where those two doors meet. It is not a guide to reading the columns of a rate card — that belongs to the sibling page on reading one rate card — and it is not the arithmetic of a single invoice, which belongs to the arithmetic of a single bill. The question here is narrower: for one vendor's models, what does each entry point charge for, where do the two prices cross, and why does a cache write reprice a workload in a way a headline rate hides? The full taxonomy of billing mechanisms is the centre page's job — the taxonomy of AI billing mechanisms — and this page takes it as given.

Claude pricing: one model family, two ways to buy it

Claude pricing resolves into two lists, and the simplest way to keep them straight is to read them as a subscription list and a metering list.

The subscription list sells seats. At the time of writing it runs from a free tier at no charge, through Pro at twenty dollars a month (seventeen a month if billed yearly, two hundred up front), to two Max tiers at one hundred and two hundred dollars a month that multiply Pro's per-session allowance by five and by twenty. Team adds seats at twenty-five dollars per seat per month on monthly billing, or twenty on annual, with a premium seat at a hundred and twenty-five that carries five times a standard seat's usage, and Enterprise is an annual seat fee plus usage that scales with the model and the task. Every number in that list buys the same thing: a right to use the models inside an allowance, described in multiples of Pro rather than in fixed message counts.

The metering list sells tokens. On the API the mid-tier Sonnet-class models sit at three dollars per million input tokens and fifteen per million output tokens, the small Haiku-tier models at one and five, and the flagship Opus tier at five and twenty-five. There is no monthly minimum and no seat: you are billed for exactly what you send and what the model writes back, separately for the two directions.

Read side by side, the two lists are not two prices for one product. A seat and a key differ in three ways that matter more than the figures. First, what is metered: a seat is not metered at all inside its allowance, so a busy day and a quiet day cost the same; a key is metered token by token, so the invoice tracks the traffic. Second, what an overrun looks like: a seat that hits its ceiling produces a blocked or slower session and an upsell, while a key that runs long simply produces a larger invoice. Third, what is visible: a seat hides the per-token columns, and the two cache columns in particular, while a key itemises input, output and both cache directions. That third difference is the one this page is really about, because the cache asymmetry is where the two entry points diverge most.

Claude Sonnet pricing: the mid-tier rate and what it buys

The Sonnet tier is the one most production teams actually pay, so it is the honest place to start. On the API a Sonnet-class model is quoted at three dollars per million input tokens and fifteen per million output tokens, the same one-to-five ratio the rest of the lineup carries. On the subscription side that model is not priced at all: it is included in the Pro and Max allowances, so its "price" in that door is a share of a monthly fee rather than a rate.

Boundaries travel with the rate. Our own price table, read for this page, carries a Sonnet entry with a two-hundred-thousand-token input ceiling and an eight-thousand-token output ceiling, and those two numbers decide more than the rate does for long-document work. A Sonnet request that reads a large retrieved document and writes a short answer is dominated by the cheap direction; a Sonnet request that reads a short instruction and writes a long report is dominated by the expensive one, and the ceiling on output is what stops the second case from running away.

The practical consequence of the two doors is a crossing point rather than a winner. At a low, steady volume, a seat is usually the cheaper door because it stops metering and its fee is bounded; the monthly fee is known in advance and does not move with a quiet week. At a spiky volume, or a programmatic one where requests arrive from a service rather than a person, the key is the honest door, because it charges for the traffic that actually happened and nothing for the traffic that did not. The mistake worth naming is buying the seat and then measuring usage as if it were the key: re-deriving a per-token rate from a fixed fee only produces a number that changes every month.

Claude API pricing: what changes when a seat becomes a key

Moving from a seat to a key is not a change of rate; it is a change of contract. Four things change at once, and each of them is a decision the subscription door had already made for you.

The first is per-model rows. A subscription makes the model choice a feature of the plan; the API makes it a row in a table, and the spread between the rows is large — a Sonnet-class row at three and fifteen sits below an Opus-class row at five and twenty-five, and above a Haiku-tier row at one and five. You now have to decide, per endpoint, which row you are buying, and a prompt that could run on either has a cost attached to the choice.

The second is the cache columns. A seat does not itemise caching; the API does, with a separate price for writing a prefix to the cache and for reading it back. That distinction has no analogue in the subscription door, and it is the reason two teams sending the same prompts can see different bills — one writes a prefix once and reads it many times, the other rewrites it on every call.

The third is the batch row. The API sells the same tokens at a discount when the work can wait in a queue, because the provider can schedule it on hardware it would otherwise idle. A subscription cannot offer that trade, because a seat is a right to use the product interactively, not a queue position.

The fourth is where the request may travel. A key is a credential you can point at a different routing layer; a seat is not. That is the whole reason reserved-capacity arrangements exist on the API side — the cloud platforms sell a committed tier that behaves like a seat but is measured like an API, so a team can buy the predictability of a subscription without giving up the meter (Azure OpenAI's provisioned deployments). The same open routing is why a genuinely free tier is credible only on the API side, where the constraint is a quota rather than a price (free-tier LLM APIs).

Anthropic API cost: where the money actually lands

If you metered the API for a month, the bill would land in five places, and knowing which place is large is the whole game. The largest single line for most generative workloads is output, priced at about five times input across the current lineup. The next is base input, the tokens read but neither cached nor reused. Then the two cache lines, which are small individually and dominate for a specific shape of workload. Last is the batch discount, which is not a line at all but a multiplier applied to the others.

Two workload shapes make that ordering concrete. A generator — an assistant that reads a short instruction and writes a page — spends almost everything in the output column, and the input rate it was quoted barely moves the total. A retriever — a service that re-sends the same long system prompt and retrieved context on every call — spends almost everything in the input columns, and its fate is decided by how much of that prefix is cached and how often the cache is read back. The two shapes can consume the same number of tokens and still have completely different bills, because they consume different columns.

Where the columns sit is also partly a vendor policy, which is why the same workload can look cheap at one provider and expensive at another without any rate being "wrong". A provider can lean on a low output rate for long-form work, so a generation-heavy job is cheapest there (DeepSeek's low per-token rates); a provider running on specialised inference hardware can push the raw per-token rate down on the models it serves (Groq's per-token rates). Both of those are still per-token doors. Neither replaces the seat-versus-key decision; they move the numbers inside the key door.

Prompt caching cost: the write-versus-read asymmetry

Prompt caching is the part of Claude pricing that most repays being read carefully, because it is deliberately asymmetric and the asymmetry runs the "wrong" way on the first call. Writing a prefix to the cache is billed above the base input rate: a five-minute write at one and a quarter times base input, and a one-hour write at twice base input. Reading that prefix back is billed at roughly a tenth of the base input rate. So a cache write is a negative discount — the first call that caches a prefix costs more than the same call without caching — and a cache read is a positive one.

Two consequences follow, and both change what "a cheaper workload" means. The first is that caching is a bet on reuse, not a saving in itself. If a prefix is written once and never read again, the cache write has cost more than the uncached input would have, and the workload is strictly worse off. If the prefix is written once and read back once, the pair of calls is already ahead, because the write premium is smaller than the read discount at a reasonable reuse count. If the prefix is written with the one-hour lifetime instead of the five-minute one, the write premium doubles, and the prefix now has to be reused more than once before the cache is ahead of simply re-sending it.

The second consequence is that the cache only ever touches the input side of the ledger. A token the model has not generated yet cannot have been cached, so caching lowers the price of context and leaves the price of output exactly where it was. That is why a caching strategy is a plan for keeping a stable prefix stable, not a plan to make a workload cheap in general: a prompt dominated by the model's own previous output gains nothing from the cache, while one dominated by a fixed document gains nearly everything.

The lifetime is the third moving part, and it is where a cache bet quietly fails. The cache entry expires — minutes by default, an hour at the higher write price — and a read after the lifetime is a miss, which means a fresh write at the premium and a bill that looks like the cache was never a saving at all. A prefix that is reused continuously stays warm; a prefix reused once an hour against a five-minute lifetime is paying the write premium forever and reading at full input price. Different providers put that same asymmetry in different places — some absorb the write and discount only the read, some make caching automatic rather than declared — so the cache columns are exactly where two per-token cards stop being comparable on the headline rate alone (OpenAI's automatic prefix caching).

Output token cost: the column caching cannot touch

Output token cost is the column a workload feels most, and it is the one no cache can reduce. Every model in the Claude family prices output at roughly five times its input, which is a structural ratio rather than a policy: reading a prompt is a parallel operation, while generating is sequential, one token at a time. The ratio is why a service that answers in one sentence is cheaper than a service that answers in one page even when both read exactly the same thing, and it is why the first lever on a generative bill is almost always answer length rather than prompt length.

The seat-versus-key interaction shows up here in an unusual way. Inside a subscription allowance, output is not separately priced: a longer answer does not cost more dollars, it consumes more of the same per-session allowance, and the ceiling shows up as a slower or blocked session rather than as a line item. On the API, output is a column with a rate, and a verbose model is more expensive by construction. A team moving from a seat to a key therefore changes the meaning of "long answer" from a usage habit into a bill, and the change is larger than the rate alone suggests, because the output rate is the expensive one.

Because caching cannot reach output, the two asymmetries compound rather than cancel. A workload that re-reads a stable context benefits twice: the cache lowers its input cost, and its short answers keep the output cost small. A workload with a fresh context on every call and long answers gains nothing from the cache and pays the full output rate on every token it writes.

Cost per token: the two Claude rows in our own price table

A cost per token is only as honest as the table behind it, so it is worth putting one in full — the one we actually run. Our budget module keeps a local JSON file that holds four values per model: input cost per token, output cost per token, maximum input tokens and maximum output tokens. The unit is one token, never "per one thousand", so there is no hidden division between the stored rate and the arithmetic that consumes it. The two Claude rows at the revision this page was read from are:

Model id Input (USD / token) Output (USD / token) Max input Max output
claude-3-5-sonnet-20241022 3.0e-6 1.5e-5 200,000 8,192
claude-3-haiku-20240307 2.5e-7 1.25e-6 200,000 4,096

Read in per-million terms, those two rows are the same list the vendor publishes: three dollars in and fifteen out for the Sonnet-class row, a quarter and one and a quarter for the small one. The table does not reproduce the subscription list, and that absence is the point — a seat is not a per-token price, so there is nothing for a per-token table to hold.

Three properties of that table are worth copying whatever the numbers are on the day you read this. First, the unit is deliberate: storing a per-token fraction removes an implicit thousand from every calculation and makes a wrong rate a wrong number rather than a wrong scale. Second, the tokenizer is chosen per model family — the counter behind this table uses one encoding for the GPT-4o and DeepSeek families and falls back to another for everything else — which is why a single character count is not comparable across those rows; the counting question itself is its own page, and it is the metre behind these rates rather than the rates themselves. Third, the table is a snapshot, not a quote: it is a local, diffable file that changes in a commit, which means a rate change is a reviewable diff and also means the table is only as current as its last edit. Our own rule applies to ourselves — the figures above carry the date they were read, and a budget built on them should re-read the file rather than trust this page.

How SmartGate fits

SmartGate does not sell Claude tokens, so it does not fit the per-token door at all, and it does not sell seats either. It is the layer in front of the tool traffic that fills an agent's context before a model is ever asked to complete anything, and it prices a platform rather than a rate:

Plan Monthly token cap MCP requests/min per key Audit-log retention Max keys per team
Free 2M 120 7 days 2
Pro 20M 300 30 days 10
Teams 100M 600 90 days 30
Enterprise 200M+ 1200 180 days 9999

Those columns are not rates — they are ceilings and retention windows that decide what a workload is allowed to do, and they are read against the two doors above as constraints rather than as prices. The billing model is the part worth stating plainly, because it is where a platform avoids behaving like a markup: pay for the platform, and share only when it saves you something — the share begins after a measured saving threshold, so the fee does not grow with every call the way a per-token markup would. If your cache-and-context strategy is the reason the model bill fell, that is the saving the model shares. The caps, per-key rates, retention windows and key counts move by tier, so the pricing page is the authoritative table and this page is not.

If you want to see what the platform layer costs before you decide between a seat and a key, start free and watch one agent run end to end; the token caps on the free tier are enough to measure a real workload rather than a guess.

Frequently Asked Questions

Is the Claude API cheaper than a Claude subscription?

Neither is cheaper in the abstract; they charge different things. The API bills per token, so a small or spiky workload pays only for what it uses, while a subscription charges a flat monthly fee for an allowance that does not move with usage. A seat is cheaper once steady usage is large enough to consume the allowance; a key is cheaper when usage is small, uneven or programmatic. Find the crossing point from your own usage history rather than from either list.

Why does a cache write cost more than a normal input token?

Because the provider is doing extra work: it is computing and storing a prefix so that a later call can skip it. That upfront cost is charged above the base input rate — a five-minute write at one and a quarter times input, a one-hour write at twice it — and the discount only arrives later, when the prefix is read back at about a tenth of the input rate. A cache write is a small bet that the prefix will be reused, and it loses that bet if the prefix is never read again before the cache expires.

Does caching reduce the cost of a long answer?

No. Caching only covers input, because a token the model has not generated yet cannot have been cached. Output stays at its own rate — about five times input across the lineup — on every call. Caching makes repeated context cheaper, which is a different thing from making a long answer cheaper.

Why is output priced at five times input?

Reading a prompt processes the whole input in one parallel pass, so the marginal cost per input token is low; generating is sequential, emitting one token, appending it, and running again, so each output token carries a higher marginal cost. The ratio is structural rather than a vendor's choice, which is why it holds across the model range and why the length of an answer drives a generative bill more than the length of the prompt does.

Can I use the API and a subscription at the same time?

Yes, and many teams do: a seat for interactive work by people, an API key for the service traffic the seat cannot meter. They are separate contracts with separate billing, so the two do not subtract from each other, and the reason to track both is reporting rather than savings.

Limitations

This page is about the interaction between two entry points for one vendor's models, and it deliberately copies no third-party figure it could not date. The subscription figures and the API rates quoted above were read from the vendor's own pages on 2026-10-08, and both lists move; the numbers are here to show the shape of the two doors, not to stand in for the vendor's current table, which always wins over any summary including this one.

It also stops at the boundary of money. It does not read the columns of a rate card, does not compute a single request's cost, does not explain how usage is counted and attributed, and offers no savings checklist or quota design — those are separate pages in this cluster, and each owns its own angle. A reader who needs an actual figure should take the door they chose here to the vendor's own card and re-check it on the date they plan to spend against it.

Finally, the plan limits and the two Claude rows are our own readings from the product source and a local price table, taken on 2026-10-08. They are operational limits and stored rates rather than a published quote, so they are exactly as current as the files they were read from; a budget built on them should re-read those files rather than trust a copy of them here.

Sources

  • Anthropic Claude API pricing, for the per-million input and output rates and the model tiers: https://docs.anthropic.com/en/about-claude/pricing
  • Anthropic prompt-caching documentation, for the cache-write multipliers (1.25x for the five-minute lifetime, 2x for the one-hour lifetime) and the cache-read rate: https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching
  • Claude subscription plans, for the Pro, Max and Team seat prices: https://claude.com/pricing
  • Claude help centre, for the Pro plan description and annual billing: https://support.claude.com/en/articles/8325606-what-is-the-pro-plan
  • Product behaviour and the plan table: read read-only from the product source at the revision recorded in this project's pipeline_results.json, with the plan figures re-verified against the live /pricing page on 2026-10-08.
  • Demand and section keywords: this project's own measured pool, recorded in search_volume.json and research_brief.md.
  • The price-table section is our own implementation, read on 2026-10-08 from backend/smartgate/modules/budget_guard/model_prices.json with backend/smartgate/modules/budget_guard/algorithm.py (origin/main): the four values held per model, the two Claude rows quoted in that section, the per-token unit, and the per-family tokenizer choice.

Method note

This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 1 of 7 sections for this lane (0 abstentions, 6 no-slice verdicts): the cost per token section resolved under rule A L1 to a per-token helper in our budget module, and the other six section keywords came back no-slice, which is expected for a vocabulary of vendor and pricing words. This batch ships as zero-code pages, so no window is cut even for the one section that pinned; its substance is the four values per model in the price-table section above, stated from the files rather than quoted as an excerpt.

Product claims were read from the product source at the revision the slice run recorded in this project's pipeline_results.json, read-only, and the plan figures were re-verified against the live pricing page on 2026-10-08. The section keyword quoted above each heading comes from this project's own paid measurement run, not from a third-party tool. No code, batch fingerprints, auction data or internal hosts appear in the text, so there is nothing here that has to be asserted verbatim.