SmartGateSmartGate

AI Token Cost: From a Price per Million to a Real Bill

AI token cost is two things multiplied together — a unit rate quoted per million tokens, and a billing rule that decides which tokens are counted at which rate. Input, output and cached-input tokens carry different prices, so the same prompt can cost several times more or less depending on what is cached, what is batched, and which model answers it.

Short answer: AI token cost is two things multiplied together — a unit rate quoted per million tokens, and a billing rule that decides which tokens are counted at which rate. Input, output and cached-input tokens carry different prices, so the same prompt can cost several times more or less depending on what is cached, what is batched, and which model answers it. This page is the accounting: how a per-million quote becomes a per-request and per-session number, and why one unchanged prompt produces a different bill under a different route.

Key takeaways

  • A rate card is three numbers, not one. Uncached input, cached input and output are priced separately; output is typically five to six times the input rate, so long answers dominate a bill.
  • Convert a per-million quote by dividing by a million. Multiply that unit rate by the counted tokens per call; the answer is a fraction of a cent, so round only at display.
  • Caching changes the rate, not the request. A repeated prefix bills at roughly 10% of the input rate on both OpenAI and Anthropic cards, which makes prompt structure a pricing decision.
  • The same prompt bills differently under routing. A cheaper model, a batch tier or a cache hit changes the counted token mix — the text did not change.
  • Caching is not free money. A cache write carries a premium, and a cache that never hits is a net loss, so the win is measured, not assumed.
  • Measure the inputs before you argue about the rate. Token counts, cached share and model mix per call are the numbers that move a bill, and they are what an LLM observability tools stack is bought to show.

ai token cost: the meter, and the two halves of it

Every AI invoice multiplies two quantities that are easy to confuse: the unit rate on the provider's card, and the counted tokens the meter recorded. Teams argue about the first and are surprised by the second, which is why a period of falling list prices has not produced falling bills. If the rate halves but the counted input triples — because the agent re-sends a growing transcript and a large tool surface every turn — the monthly number still rises. The rate you can look up; the count you have to instrument.

The meter is narrower than the bill, and it prices at least three classes of token:

Token class What it is How it is priced
Uncached input Everything you send that the provider has not seen recently: the prompt, the conversation so far, tool schemas, pasted documents the full input rate
Cached input A prefix the provider already processed and stored, billed on a later read about one tenth of the input rate on the current cards
Output The tokens the model generates the output rate, usually five to six times input

Two consequences follow. A card that quotes a single "price per token" is quoting half a price; and the honest unit of comparison between two providers is not the headline input rate but the blended cost at your own token mix — how much of your input is a stable, cacheable prefix, and how long a typical answer runs.

The other half of the meter is time. A token cost is not one number per month; it is a series of counts attributed to the operation that produced them, and reading it as a series is what turns "tokens got expensive" into a target. Which call, which tool, which route, on which day — that is a record-keeping problem, and it is the layer this page assumes before it does any arithmetic.

token cost: from a per-million quote to one request

The conversion from a rate card to a request is one line of arithmetic, and almost every estimate gets it wrong the same way — by treating a million-token price as a per-token price.

Cost per request = (uncached input / 1,000,000 x input rate) + (cached input / 1,000,000 x cached rate) + (output / 1,000,000 x output rate).

Worked on the current OpenAI list card, using GPT-5.6 Terra at $2 per million input, $0.20 per million cached input and $12 per million output (OpenAI list price, checked 2026-10-03): a support assistant sends a 2,000-token prompt, of which 1,600 is a cached system prompt, and returns 500 tokens. The uncached 400 input tokens cost $0.0008, the 1,600 cached tokens cost $0.00032, and the 500 output tokens cost $0.006 — about $0.0071 per call. Uncached, the same call costs $0.004 of input plus $0.006 of output, or $0.01: caching the prefix removed roughly 29%, and output tokens are still more than half the bill.

That is the surprise. Output is six times the input price on those rates, so a workload that writes a lot costs more than one that reads a lot, even with the same prompt. The way to keep the arithmetic honest is to read token cost as a dated series rather than a single figure, which is what the pinned excerpt does: it merges a daily output-token series with a per-operation savings series into one row per date, so cost is attributed by day and by operation instead of collapsed into a monthly total.

# lib/dashboard/reports-trend-metrics.ts — source lines 124–164 (token cost as a dated series)
function buildTokenCostSeries(
  dailyTokens: ReportsTrendDataSlice["dailyTokens"],
  costSavings: ReportsTrendDataSlice["costSavings"],
) {
  const byDate = new Map<
    string,
    {
      date: string;
      output: number;
      compress: number;
      fetch: number;
      search: number;
    }
  >();
  for (const row of dailyTokens) {
    byDate.set(row.date, {
      date: row.date,
      output: row.output,
      compress: 0,
      fetch: 0,
      search: 0,
    });
  }
  for (const row of costSavings) {
    const prev = byDate.get(row.date);
    if (prev) {
      prev.compress = row.compress;
      prev.fetch = row.fetch;
      prev.search = row.search;
    } else {
      byDate.set(row.date, {
        date: row.date,
        output: 0,
        compress: row.compress,
        fetch: row.fetch,
        search: row.search,
      });
    }
  }
  return [...byDate.values()].sort((a, b) => a.date.localeCompare(b.date));
}

The shape of that function is the shape of the problem. There is no single "cost" field; there is an output count and three separate operation counts (compression, fetch, search) on a date key, sorted before it is read. A dashboard plotting one blended number can tell you the month was expensive; a series keyed by date and operation can tell you which of the four is growing. The conversion above is cheap; the attribution it depends on is the hard part.

cost per token: the unit rate and how to convert it

The unit rate is where the per-million quote becomes small enough to multiply. At $2 per million, one input token costs $0.000002; at $12 per million, one output token costs $0.000012. Those are the figures the code multiplies, and it is worth reading for three choices it makes explicitly.

First, input and output are separate terms with separate prices — there is no single blended rate. Second, a caller may pass a custom price table, which is how a negotiated or private deployment keeps the same arithmetic. Third, the result is rounded to six decimal places on the way out; that is the tell that a per-token cost is routinely smaller than a cent, so precision must live in the accumulation and be trimmed only once the numbers are large enough to display.

# backend/smartgate/modules/budget_guard/algorithm.py — source lines 111–138 (per-token unit cost)
def cost_per_token(
    model: str = "",
    prompt_tokens: int = 0,
    completion_tokens: int = 0,
    custom_cost_per_token: Optional[Dict[str, float]] = None,
) -> Dict[str, float]:
    """提取自 litellm/cost_calculator.py cost_per_token() — 核心计算逻辑。"""
    # Custom pricing
    if custom_cost_per_token is not None:
        input_cost = custom_cost_per_token.get("input_cost_per_token", 0) * prompt_tokens
        output_cost = custom_cost_per_token.get("output_cost_per_token", 0) * completion_tokens
        return {"input_cost": input_cost, "output_cost": output_cost, "total_cost": input_cost + output_cost}

    # Lookup model prices
    prices = _get_model_prices()
    model_info = prices.get(model, {})

    input_price = model_info.get("input_cost_per_token", 0)
    output_price = model_info.get("output_cost_per_token", 0)

    input_cost = prompt_tokens * input_price
    output_cost = completion_tokens * output_price

    return {
        "input_cost": round(input_cost, 6),
        "output_cost": round(output_cost, 6),
        "total_cost": round(input_cost + output_cost, 6),
    }

Read as accounting, that function says what a rate card cannot: the per-token rate is not the cost of anything by itself. It is a coefficient, and the cost is whatever token mix you multiply it by. A model with a lower input rate but a higher output-to-input ratio can be the more expensive choice for a chat workload and the cheaper one for short classification, on the same coefficient table. That is why the useful comparison later is not "which model is cheaper" but "which is cheaper at this mix", and why the number teams argue about at month end is a mix, not a rate.

The discipline that follows is small. Record the token counts per call (uncached input, cached input, output) next to the model, and recompute the bill from the rate card instead of trusting one aggregated counter. When a bill moves, the record says whether the rate moved or the mix did. Taking those counts in the first place — tokenizer versions, framing tokens, cache accounting, provider usage against a local estimate — is the ai token usage problem, and this page assumes its output. Turning those per-call counts into a dated, per-tool spend view is the job of the execution-cost breakdown, and the reductions that change the mix in your favour are collected as token optimization techniques.

llm api pricing: the three columns of a rate card

LLM API pricing has converged on a shape, and once you can read the shape you can compare cards that quote different headline numbers. A modern card is a matrix of multipliers on a base input rate:

Row on the card What it multiplies Typical value on the current cards
Standard input the base rate the number in the headline
Cached input (read) the base input rate about 0.1x
Cache write the base input rate 1.25x for a 5-minute cache, 2x for 1 hour
Output the base output rate 5x to 6x the input rate
Batch both input and output 0.5x, for work that can wait up to 24 hours
Long context input and output 2x input, 1.5x output, above a context threshold

The cache row decides real bills, and it comes in two halves that are easy to conflate. A cache read is cheap — roughly a tenth of input — because those tokens are not re-processed. A cache write is not; you pay a premium once to store the prefix. So caching pays off on repeated prefixes (a system prompt, a tool schema, a fixed document) and loses on a prefix used once. The right question is never "should we cache?" but "how many reads does this prefix get before it changes?"

Two more rows shape a realistic estimate. Batch halves both directions for asynchronous work and stacks with caching, so a batch job reusing a cached prefix can land well below the standard rate. Long context raises the input rate past a threshold, which is why a larger context window can cost more per token, not less. None of these rows changes the prompt; they change which rate the meter applies to it. Where the base rate itself comes from — GPU hours, utilization and the prefill/decode split — is the inference cost layer below the card, and this page treats the card as its input.

llm pricing comparison: reading two rate cards without guessing

Comparing providers on headline input price is the most common mistake here, because input is not where most of the money goes. The two columns that decide a production bill are the cached-input rate (how cheaply a repeated prefix re-reads) and the output rate (the cost of each generated token); their ratio tells you which workload shape a model is priced for. Reading a whole catalogue on those two columns, before any arithmetic, is ai model pricing; the table below is that reading applied to this page's own examples.

Model (list price, USD per 1M tokens) Input Cached input Output Output : input
OpenAI GPT-5.6 Sol 4.00 0.40 20.00 5x
OpenAI GPT-5.6 Terra 2.00 0.20 12.00 6x
OpenAI GPT-5.6 Luna 0.20 0.02 1.20 6x
Claude Opus 5 5.00 0.50 25.00 5x
Claude Sonnet 5 2.00 0.20 10.00 5x
Claude Haiku 4.5 1.00 0.10 5.00 5x

Figures are the providers' published list prices, standard tier, checked 2026-10-03 (OpenAI GPT-5.6 family from the OpenAI pricing page; Claude values from the Claude API pricing table). Three readings come from the columns, not from any cell. First, the cached-input column is almost exactly proportional to input everywhere (about 10%), so on these cards caching does not differentiate the providers — the base rates do. Second, the output column is where the families separate: two models can share a $2 input rate and differ by 20% on output, which flips the cheaper choice for a generation-heavy workload. Third, the honest comparison for one application is a weighted blend: multiply each column by your own shares of uncached input, cached input and output, then compare totals — not the headline rows.

Two caveats belong with any such table. List prices move and promotions expire, so a comparison means nothing without a checked date and a link to the provider's own card as the authority. And a list price is not a contract: a smaller model with a shorter context window can cost more in practice if it forces extra calls or a longer answer to compensate.

openai api cost: a worked estimate from the published list

OpenAI bills per token against a rate card, with no platform fee and no monthly minimum — the meter is the product. For the GPT-5.6 family the structure is the one above: separate input, cached-input and output rates; a cache-write premium of 1.25x input on the models that support it; an automatic 0.1x cached-read rate for a repeated prefix; and a Batch tier at half price for jobs that finish within 24 hours (OpenAI list price, checked 2026-10-03). Output is about six times input across the family, so an estimate that counts only prompt size will undershoot.

A monthly estimate is the per-request arithmetic scaled up. At about $0.0071 per cached support call, 20,000 calls a day is roughly $142 a day, about $4,270 a month. The same workload with no caching runs at about $0.01 a call, roughly $6,000 a month — a gap that is entirely the 1,600-token prefix, the single largest lever in the estimate. That is the shape of an OpenAI estimate worth trusting: a call count, a per-call token mix, and a rate per class, multiplied.

The other rows quietly move an estimate. Batch halves both directions and stacks with caching, so a nightly backfill on Batch with a cached prefix costs a fraction of the same job run live. Long-context requests bill input at 2x and output at 1.5x past the threshold, so a large-context request is a different unit price, not the same one. Cache writes carry the 1.25x premium, so a prefix that changes every request is a worse deal than no caching. An estimate that names the mix, the cache behaviour, the batch tier and the context size is one you can defend; an estimate that names only the model is a guess with a decimal point.

anthropic api cost: the same arithmetic, different numbers

Anthropic's Claude API uses the same three-column arithmetic with one structural difference: caching is an explicit, priced operation with a duration, not an implicit automatic discount. A Claude cache write costs 1.25x the base input rate for a 5-minute cache and 2x for an hour; a cache read then costs 0.1x the base input rate. The provider's own guidance follows: a 5-minute cache pays for itself after one read, a 1-hour cache after two, because the write premium is only recovered by cheap reads (Claude API pricing, checked 2026-10-03).

Worked on Claude Sonnet 5 at $2 per million input, $0.20 per million cached input and $10 per million output: an agent session of 50,000 input tokens, of which 40,000 is a stable cached prefix, plus 3,000 output tokens, costs $0.008 of cached reads, $0.02 of uncached input and $0.03 of output — about $0.058 per session. Without caching, the same 50,000 input tokens cost $0.10 and the session lands at $0.13. The 40,000-token prefix would be written once for $0.10 (40,000 tokens at 2x base input, 5-minute duration); after two re-reads that cost is behind you and every later read is cheap.

The duration has no OpenAI equivalent, and it changes the answer to "should we cache this?". A prefix reused within the cache's lifetime earns the discount; one used once an hour against a 5-minute cache does not, because each request pays a write premium and never reads. So an Anthropic estimate needs one extra input — the reuse interval of each prefix — before the number means anything. With it, the arithmetic is identical to OpenAI's. Cross-provider, the move is to keep the formula and swap the rate table.

How SmartGate's price sits next to model token cost

The model bill above prices the completion your application asks for. A large part of a real AI bill is not that call — it is the tool traffic that fills the context before the completion is requested, and that traffic is governed a layer earlier. SmartGate is priced as a platform for that layer, and its numbers are operational limits rather than per-token rates:

Plan Monthly token cap MCP requests/min per key Audit-log retention Max keys per team
Free 2M 120 7 days 2
Pro 20M 300 30 days 10
Teams 100M 600 90 days 30
Enterprise 200M+ 1200 180 days 9999

The billing model is the differentiator worth stating plainly: pay for the platform, and share only when it saves you something — the share starts after a measured savings threshold, and the Pro total is capped, so the fee does not scale with every tool call the way a per-token proxy would. Read against the rate cards above, the two layers compose: the gateway decides which tool calls happen and caps the token traffic they generate, and the provider's card prices the completion that follows. The caps, per-key rates, retention and key counts are operational limits that change with the plan; the pricing page is the authoritative table, and a ceiling that outlives a budget review is the FinOps lead's view rather than a rate. Where the enforcement mechanics live — reading a team ceiling, counting a call, refusing at the check — is worked through on enforcing a token quota per team.

How to get started

  1. Pick one workload and one model. Compute its per-call cost from the provider's own card with the three-column formula — input, cached input, output — and write the figure down. If you cannot name the token mix, you are not ready to compare providers.
  2. Measure the mix, not the total. Record uncached input, cached input and output per request, plus the model that answered. The blended number is an output of the record, never an input to it.
  3. Test caching on a real prefix. Cache the system prompt, the tool schema or a fixed document, and compare cached and uncached cost over a week. A prefix read often pays; one that changes every request does not.
  4. Compare the cards at your own mix. Weight each provider's three rates by the shares you measured and compare totals, with the checked date beside the numbers. Start on the free platform tier to see the tool-traffic half of the bill on real data with start free, then confirm the current plan limits against your own workload.

Frequently Asked Questions

Why did my bill rise when the per-token price fell?

Because the bill multiplies a rate by a count, and the count moved more than the rate. A growing transcript, a larger tool surface or a retry loop all raise counted input tokens per turn, and output tokens are the expensive class, so a lower input rate on a longer workload still produces a bigger number. Compare the counted mix period over period, not the headline rate.

Is input or output more expensive?

Output, by five to six times on the current cards. Terra lists output at six times input and Claude Sonnet 5 at five times. That ratio decides which workloads are expensive: a summarisation or code-generation job that writes a lot costs more than a retrieval or classification job that mostly reads, even with a similar prompt.

Does prompt caching always save money?

No. A cache read is cheap, about a tenth of input, but a cache write carries a premium of 1.25x for five minutes or 2x for an hour on Claude, so caching only pays when a prefix is read often enough to recover the write. A prefix that changes every request, or is reused more slowly than the cache lifetime, is a net loss.

How do I estimate the cost of one agent session?

Add three numbers: uncached input tokens across all turns, cached input tokens across all turns, and total output tokens. Multiply each by its unit rate (list price divided by one million) and sum. A 60,000-input and 8,000-output session on Terra with a 50,000-token cached prefix is about $0.13, and the cached share is most of the difference from the uncached $0.22.

Do list prices apply to batch and long-context work?

Neither uses the standard rate. Batch halves input and output for asynchronous jobs, and long-context requests above the threshold bill input at 2x and output at 1.5x. Because batch and caching stack, an overnight job reusing a cached prefix can run several times below the standard rate — a different unit price, not a rounding difference.

How does SmartGate's price relate to model token cost?

They are two layers. The provider's card prices the completion your app asks for; SmartGate governs the tool traffic that fills the context before that request, meters it against a team token cap, and charges a platform fee rather than a per-token markup, taking a share only once it has measurably saved you money. The two bills add rather than replace each other.

Limitations

This page explains a pricing model, and a model is only as good as the numbers fed into it. The list prices quoted were checked on 2026-10-03 and are the providers' published figures, not a commitment by us or by them; rate cards change, promotions expire, and negotiated or regional rates differ from the public table. Any estimate built here should be recomputed against the provider's own card before it goes into a budget.

The arithmetic also assumes a clean token mix. A request can carry imagery, audio, tool calls and storage line items priced on separate rows and not modelled here, so a real invoice can exceed a token-only estimate. And the cached-input discount is a rate, not a guarantee: providers decide what is cacheable and for how long, and a prefix that misses the cache bills at the full input rate however it was written.

Finally, this page deliberately does not cover the savings checklist, quota design, or monitoring. Those are separate pages in this cluster, linked above, and each owns its angle; one page claiming all four would answer none of them well.

Sources

Method note

This page carries two code excerpts out of seven sections, and that split is the honest result of the slice matcher, not an editorial choice. pipeline.alethix run pinned 2 of 7 sections for this page, with 0 abstention(s) and 5 no-slice verdict(s): rule A found a unique symbol for the token cost section (rule A level 3) and for the cost per token section (rule A level 1), both confirmed by a server-side proof call, so those two sections carry a fenced block cut verbatim from the slice body. The other five section keywords returned either no local match or a set of equally plausible candidates that could not be uniquely resolved, so those sections were written from external sources — the providers' own pricing documentation, cited above with a checked date. A generic symbol pinned into those sections would have given the page the look of verified code with none of the substance, and the house rule is that an unpinned section is sourced, never invented.

Product claims were read read-only from the product source at the revision the slice run recorded in this project's pipeline_results.json, and the plan figures were re-verified against the live pricing page on 2026-10-03. Numeric list prices in the pricing sections are the providers' published figures with their own checked date, not our commitments. The section keyword quoted above each heading comes from this project's own paid measurement run. No code is transcribed by hand, no batch fingerprints, auction data or internal hosts appear in the text, and every fenced block above was re-asserted byte-for-byte against the slice body before publication.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)
1 token cost buildTokenCostSeries lib/dashboard/reports-trend-metrics.ts 124–164 rule A L3 → slot-proof d265de9f668c
2 cost per token cost_per_token backend/smartgate/modules/budget_guard/algorithm.py 111–138 rule A L1 → slot-proof 2712bf90f44f