AI Token Usage: How the Number Is Counted and Attributed
AI token usage is the count of tokens a request moves through a model, split by what they are (input, output, and cached reads and writes) and attributed to whatever caused them (a caller, a team, a tool, a task). The arithmetic is trivial; the hard parts are deciding what belongs in the prompt, accepting that two tokenizers count the same text differently, and reconciling your own count against…
Short answer: AI token usage is the count of tokens a request moves through a model, split by what they are (input, output, and cached reads and writes) and attributed to whatever caused them (a caller, a team, a tool, a task). The arithmetic is trivial; the hard parts are deciding what belongs in the prompt, accepting that two tokenizers count the same text differently, and reconciling your own count against the provider-reported one when the two disagree.
Key takeaways
- A count is a quantity, not a price. Usage is what the meter records; cost is that quantity times a rate, which is a separate subject.
- What you type is a fraction of what is billed. The system prompt, the tool schemas, the conversation history and any cached prefix are all part of the input count.
- The tokenizer decides the number. One string maps to different token counts under different vocabularies, so a pre-flight estimate and the provider's count will differ.
- A count without an owner cannot be acted on. Attribution — key, team, tool, task — is what turns a monthly total into a decision.
- Reconcile before you trust it. Sample your own count against the provider's reported usage, measure the gap, and only then budget or bill on the number.
ai token usage: what the number actually counts
The phrase sounds like one number, but a usage record is at least three. Input tokens are everything the model reads before it writes: the system prompt, the user messages, every earlier turn that was re-sent, the schemas of the tools the model may call, and any document, file or search result that was pasted into the context. Output tokens are what the model generates in reply, including the hidden reasoning a reasoning model may bill as output. Cached tokens are a third pair that providers report separately — a cache read, where a prefix the provider already stored is billed again cheaply, and a cache write, where that prefix is stored in the first place.
The gap between what a person types and what gets counted is the whole reason this is worth a page. A short question on top of a fixed 1,500-token system prompt and a 2,000-token tool surface is a 3,500-token request, not a 20-token one. The typed words did not change; the prompt did. Two things follow. First, a request's count is dominated by the scaffold around the question, not by the question. Second, the scaffold is the part that repeats, which is exactly why it is worth caching and worth measuring.
It is worth being precise about the boundary, because usage and cost are neighbours that get confused. This page counts; the unit-price model is the subject of its sibling AI token cost, which turns a counted token mix into dollars. If you read one page for the arithmetic and the other for the count, the pair covers the whole invoice. Reading only one of them is how teams end up either measuring carefully but unable to price, or pricing carefully against a count nobody checked.
A second boundary matters as much: usage is not the same as a limit. A count says what happened; a cap says what is allowed. The sections below separate the two by design, because conflating them is how a metering job quietly becomes an enforcement job and starts refusing traffic it was only meant to describe.
token usage tracking: from a request to a row you can query
A count that is not attached to anything is a monthly total, and a monthly total cannot answer the only questions people actually ask: which team, which tool, which task, on which day. Tracking is the discipline of turning each call into a row with enough fields to route a question to an answer.
The fields that earn their place are small in number. A timestamp and a model name place the call in time and on a rate card. An identity — an API key identifier, and through it a team and an owner — says who pays. A route or endpoint says which surface the call came through. A tool or function name says what the model spent the tokens on, which is the difference between "research is expensive" and "one badly bounded search tool is expensive". A task or trace identifier stitches the individual calls of one agent run back into a single logical unit, because a fifty-step run is one thing a reader cares about and fifty rows in a table. Finally, the counts themselves: input, output, cached read, cached write, and the outcome, so a retry that failed is visible next to the retry that succeeded.
Two design choices decide whether the rows are useful or decorative. The first is where the count comes from. Counting on the way out (a tokenizer run before you send) tells you what you intended to spend and is available even when the call fails; counting from the provider's response tells you what was actually billed and is the number to trust. The honest record keeps both when it can, and marks which is which, so a gap between them is visible rather than averaged away. The second is where the row is written. A record that lives only inside the application that made the call dies with that application and is invisible to a gateway, a finance system or an incident review; a record in a store that several systems read is what makes the same data reusable. That store and its query surface are the job of an LLM observability tools stack, which is the layer this page assumes rather than re-specifies.
The practical test for a tracking schema is whether it can answer three questions without a spreadsheet: how much did team X spend this month, which tool moved the most tokens last week, and what did this specific task cost end to end. If any of the three needs a join nobody has written, the schema is missing a field, not the query.
token metering: why the same text counts differently
Metering is the act of turning traffic into countable units, and its central fact is uncomfortable: the unit is not stable across vendors. A tokenizer is a model-specific mapping from text to a vocabulary of sub-word pieces, and every family maps a string differently. The English approximation everyone quotes — roughly four characters, or about three quarters of a word, per token — is a rule of thumb for prose and a bad predictor everywhere else. Punctuation, indentation, long identifiers, hashtags, URLs and non-Latin scripts tokenize far less efficiently, which is why a page of code or a block of Chinese can cost several times the tokens of an equal-length English paragraph.
The drift has four common sources, and naming them is most of the fix. First, vocabulary version: the same provider has shipped multiple vocabularies, and a model on a newer one will count the identical string differently from an older sibling. Second, message framing: chat formats wrap each message with role and separator tokens that are billed but never typed, so a multi-turn conversation costs more than the sum of its text. Third, special and structured tokens: a tool call, a JSON schema or a structured-output mode adds control tokens on top of the visible content. Fourth, cache accounting: a served prefix appears as a cache read on a later call, so the same message is counted at a different rate and sometimes under a different line item than on the call that first sent it.
The consequence for a meter is that a self-count is an estimate, not a verdict. A pre-flight counter built on an approximate tokenizer is right enough to refuse a request that would blow a budget and wrong enough that you should never reconcile a customer's invoice against it. The discipline is to label the two numbers, keep the provider's returned usage as the source of truth when it exists, and treat the local count as a planning input. Where the counted unit meets the billed rate — how a cache read is priced against a live token, and why the same prompt bills differently on a different route — is worked through on inference cost.
mcp token usage: the schema is part of the prompt
Model Context Protocol changes the arithmetic in a way that surprises people who come from a single chat box, because on MCP the tools themselves are input. A server advertises its tools, and each advertisement is a name, a description and a JSON schema for the arguments; a client injects that catalogue into the model's context, usually on every turn of a conversation. A server with a handful of tools can add a few hundred tokens to every call before the user says anything; a client that mounts a dozen servers can spend thousands of tokens per turn describing capabilities the model may never use.
Three MCP-specific effects follow. The tool catalogue is a per-turn tax. Because the schemas ride in the prompt, their size multiplies by the number of turns, not the number of tool calls — a conversation that never invokes a tool still pays for the schemas. Tool results come back as input. When a tool returns a JSON payload, a document or a page of search results, that content enters the context and is billed as input on the next turn; an unbounded tool output is therefore a recurring cost, not a one-off. Sampling and resources add their own traffic. A server that asks the client's model to summarise something spends the caller's tokens from inside a tool, which is easy to miss because no application code appears to have made a model call.
The practical response is to meter at the tool boundary, not only at the model boundary. A record that names the tool that caused a turn, and the size of its schema and its returned payload, makes the catalogue cost visible and lets a team trade a broad tool surface against a narrow one. What that trade costs, once the tokens are priced, is a model-selection question handled on AI model pricing; what it should be limited to is a quota question handled below.
What our metering counts, and how it estimates
Every count above is a choice, so here is the one we made, read on 2026-10-07 from
backend/smartgate/core/record_usage.py.
- One counter per team per month. Usage accumulates in a single Redis key,
usage_month:{team_id}:{YYYY-MM}— the same key the budget guard reads, so the number on a dashboard and the number a quota is enforced against cannot drift apart. - Four tools are billable:
compress,fetch,search,dedup. The checking tools —budget_check,budget_count,rate_limit— are explicitly skipped, which is the only sane ordering: a client must not be able to spend budget by asking how much budget it has left. - The estimate is declared, not hidden. When a module does not report its own count, the recorder
derives one: a fetch counts
len(markdown) // 4, a search counts the characters of every returned snippet// 4, compression counts the tokens that went in (origin_tokens, not the smaller output), andbudget_countreports its owntokens. // 4is a heuristic, and it is on the record. Four characters per token fits English prose and over-counts code, CJK and markup-heavy pages. The point is not precision but consistency: because the method is written down and applied identically to every caller, a reconciliation dispute is about the method rather than about a number nobody can reproduce.
usage based pricing ai: metering is the bill's input
When AI is sold by consumption rather than by seat, the meter stops being an internal dashboard and becomes the invoice line. That changes the standard the count has to meet. For internal cost reporting, a count that is directionally right and cheap to produce is enough. For customer billing, the same count must be complete (nothing spends without a row), attributable (every row has an owner), reproducible (the same traffic yields the same number on a re-run) and auditable (the customer can be shown how a total was reached). A meter that fails any of the four turns into support tickets, and a disputed invoice is far more expensive than the tokens in it.
This is where the earlier sections converge, and where a usage-based product has to be honest about its error bar. A provider that resells model calls and bills by the token is quoting counts it did not generate — it relays the model provider's reported usage, and its own margin depends on how faithfully it passes those counts through. Two failure modes are common. If the reseller counts locally with an approximate tokenizer, its counts drift from the upstream ones by a few percent that looks like a rounding error until it is multiplied across a month of traffic. If it counts only what it chose to send, it misses the tokens a retry, a fallback route or a tool call added downstream. The defensive design is to record both the count you expect and the usage the upstream reported, and to surface the difference rather than hide it. The pricing side of this exchange — how providers set the rates that turn a count into an amount — is the subject of AI model pricing, and it is kept deliberately separate from the count described here.
mcp quota: when a counted number becomes a ceiling
A quota is a count with a limit attached, and the moment you add the limit you leave metering and enter enforcement. The distinction is easy to lose and expensive to forget: a meter observes, a quota refuses. Teams that describe a quota but implement only a meter discover the difference at the invoice, because nothing stopped the traffic that produced it.
The two ceilings people call "quota" measure different things and fail differently. A request limit — calls per minute per key, say — caps the rate at which work arrives and protects a service from being hammered; it says nothing about how many tokens each call costs, so a hundred cheap calls and a hundred enormous ones look identical to it. A token budget caps the total consumption over a period and is what actually bounds spend, but it cannot react to a burst the way a rate limit can. A mature gateway enforces both, because they answer different questions: the rate limit keeps the door from being kicked in, and the budget keeps the month's total from running away. To count against a budget you need exactly the attribution the tracking section described — a token count bucketed by owner and period — which is why quota design sits downstream of metering rather than beside it.
How a gateway reads an entitlement, decides the next call, refuses at the check and warns the owner before it is late is a topic in its own right, and this page does not re-derive it. The mechanics live on enforcing a token quota per team; what this page owns is the count that such a gate compares against, and the reason a quota is only as good as the number under it.
llm token usage: reconciling self-report against provider report
Every metering system eventually faces the same moment: the number it counted and the number the provider reports do not match. The reflex is to assume one is wrong. The productive move is to treat the gap as a measurement, with a known set of causes and a procedure for narrowing it down.
Start by making the two counts comparable. Provider usage is usually returned per response as an input count and an output count, sometimes with cached reads and writes broken out separately; your own count is whatever you instrumented, possibly including a tokenizer estimate and possibly missing retries and streamed continuations. Sum both over the same window, the same keys and the same period, then compare. A small, stable percentage difference is normal and usually traces to tokenizer version or to framing tokens your local counter does not model. A large or erratic difference points at something structural: calls your meter never saw, a fallback route that bypasses your instrumentation, double counting when a retried request succeeds, or streaming paths where the final usage block was not captured.
When the provider does not report usage — some streaming APIs, many self-hosted or open-weight models, and some tool calls that run a model on the server side — estimation is the only option and should be labelled as such. The workable method is sampling: count tokens precisely on a representative sample of requests with the model's own tokenizer, compute a per-request distribution, and extrapolate the total with an explicit confidence band rather than a false-precision figure. Sample across request shapes (short chat versus long tool output) rather than across time, because the cost driver is the shape, not the hour. The reconciliation procedure is the audit trail any token bill rests on, and the difference between an estimate and a measured count is the first thing a finance review will ask about.
Where SmartGate fits
SmartGate is one place in this stack where the count and the enforcement point are the same object. Every tool call through the gateway is written as an audit row — caller, route, transport, tool, token count, latency, outcome — and the same row is what the per-key rate limit and the per-team token budget read. That is the practical answer to the attribution problem: the count and its owner are recorded together, at the moment of the call, rather than reconstructed later from a provider invoice.
The plan table sets the operational ceilings a metering practice has to live inside. Monthly token caps are 2M, 20M, 100M and 200M+; requests per minute per key are 120, 300, 600 and 1200; audit-log retention is 7, 30, 90 or 180 days; and a team can hold 2, 10, 30 or up to 9999 keys. Those rows are what a team counts against, which is why the pricing page is the authoritative table and should be read as such rather than quoted from a second-hand summary. If your attribution requirement is driven by a budget review, the FinOps lead's view of the same rows is the reading that has to survive it.
Frequently Asked Questions
Is AI token usage the same thing as cost?
No. Usage is the quantity measured, in tokens; cost is that quantity multiplied by one or more rates. A usage figure with no rate attached cannot tell you what anything costs, and a cost figure with no usage breakdown cannot tell you why it moved.
Why does my token count differ from the provider's?
The usual causes are a different tokenizer vocabulary, framing and control tokens your local counter does not model, calls your instrumentation missed, and cached reads counted at a different point than you expected. A small stable gap is normal; a large or erratic one usually means calls are going somewhere your meter cannot see.
How do you estimate usage when the provider does not report tokens?
Sample a representative set of requests, count them precisely with the model's own tokenizer, and extrapolate with an explicit confidence band. Sample across request shapes, because length and tool-output size drive the count far more than the time of day. Treat the result as an estimate and say so.
Do MCP tool schemas count as tokens on every turn?
Yes, in the common design where the client injects the tool catalogue into the context on each turn. The schemas and their descriptions are input tokens even when no tool is called, so a broad tool surface is a recurring per-turn cost rather than a one-off.
Is a request-per-minute limit the same as a token quota?
No, and that is the point of keeping both. A request limit caps how often work arrives and protects the service. A token quota caps total consumption and bounds spend. Neither substitutes for the other, because a cheap call and an enormous one look identical to a request counter.
Limitations
This page describes how a usage number is produced and attributed; it deliberately does not cover how that number is priced, which is its sibling topic, and it does not describe enforcement mechanics beyond naming where they live. The question of spending less on the same usage is a separate one, worked through on optimize ai agent execution cost. It is a metering reference, not a benchmark of any provider's accuracy, and it makes no claim that any particular gateway or library implements the record fields it recommends.
Tokenizer behaviour changes with model versions, so the qualitative statements here about drift are durable while any specific character-per-token ratio is an approximation that ages. The product plan figures were re-verified against the live pricing page on 2026-10-03 and are operational limits, not a feature comparison; a compliance or budget decision should read the current table rather than this page. This page carries no code excerpt, and the reason is recorded in the method note below: the slice matcher found no unique symbol for any of its seven sections. The consequence is honest and specific — no quoted implementation, no line-numbered claim, and no assertion about the internal count shape of any product named above.
Sources
- OpenAI's help article on counting tokens — Understanding and counting tokens, for the tokenizer-as-vocabulary model and the character-per-token approximation.
- OpenAI's prompt-caching documentation — platform.openai.com/docs/guides/prompt-caching, for cached reads and writes being reported and priced as their own token classes.
- Anthropic's token-counting documentation — docs.anthropic.com/en/docs/build-with-claude/token-counting, for pre-flight counting being an estimate against the model's own tokenizer.
- The Model Context Protocol specification, server tools — modelcontextprotocol.io/specification, for the tool listing (name, description, input schema) that a client injects into context.
- OpenTelemetry semantic conventions for generative AI — opentelemetry.io/docs/specs/semconv/gen-ai, for the attribute names a usage record can adopt instead of bespoke fields.
- Demand figures in this page are our own measurements: DataForSEO Google Ads, United States, 12-month window, measured 2026-10-03, recorded in this project's search_volume.json and research_brief.md.
- Product behaviour and the plan table: read from the product source at the revision pinned in this project's pipeline_results.json, read-only, with the plan figures re-verified against the live pricing page on 2026-10-03.
- The metering section is our own implementation, read on 2026-10-07 from
backend/smartgate/core/record_usage.py(origin/main): the per-team monthly counter key, the billable tools (compress,fetch,search,dedup), the skipped checking tools, and the character-based estimation used when a module reports no count.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice
matcher pinned 0 of 7 sections for this page (0 abstention(s),
7 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of
the seven section keywords, because this lane's vocabulary — usage, metering, tracking, token
— collides with generic counter, analytics and price-helper names across a product codebase. A
pinned generic name would have given the page the shape of a verified article with none of the
substance, so every section above but one is written from public sources. The exception is "What our
metering counts, and how it estimates" — our own usage recorder, read on 2026-10-07 from
backend/smartgate/core/record_usage.py: the monthly counter key, the billable and skipped tool sets,
and the estimation rule applied when a module reports no count of its own.
Product claims were read from the product source at the revision the slice run recorded in this project's pipeline_results.json, read-only, and the plan figures were re-verified against the live pricing page on 2026-10-03. The section keyword quoted above each heading comes from this project's own paid measurement run, not from a third-party tool. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.