Short answer: Groq pricing is a per-token rate card with a speed claim attached: you pay per million input and output tokens, but the product being sold is tokens per second from purpose-built inference hardware. Those two units only meet inside a workload. For an interactive assistant the tokens-per-second figure is a capacity number — it decides how many concurrent sessions one budget can serve — so a faster endpoint can be cheaper per session while costing exactly the same per token. For a batch job the same speed is worth almost nothing, and the asynchronous tier's halved rate is the better buy.
Key takeaways
- The rate card prices tokens; the marketing prices seconds. A per-token figure and a tokens-per-second figure are different units, and no card converts one into the other for your load.
- Speed is capacity, not a discount. At an identical price per token, a faster endpoint finishes each request sooner, so a fixed concurrency clears more requests in the same hour.
- Faster is cheaper only for interactive traffic. When a person waits on the response, latency sets how much capacity you need; when a job can wait, the batch tier's 50 percent discount beats any speed premium.
- Rate limits are the real throughput price. A per-key request cap and a separate team ceiling decide how many calls a fleet can make per minute, and it is usually the ceiling that binds.
- Do this first: label the workload interactive, batch, or long-input before comparing anything, then price only the tier that matches the label.
Most confusion about groq pricing comes from reading a token rate and a speed claim as if they were the same sentence. They are two sentences about two quantities. The rate card answers "what does a token cost"; the speed figure answers "how quickly do tokens arrive". A buyer who multiplies the first by an expected token count gets a monthly bill and believes the question is closed, when the figure that decides the experience — and often the capacity bill — was never in the multiplication at all.
This page is about that gap. The centre of this cluster maps the whole family of billing mechanisms; a sibling works the compute economics of running the silicon yourself; others handle per-model rate cards and the counting question. This page keeps to one lane inside a single per-token card: it treats speed as a pricing axis, asks how a speed tier, a batch tier and a concurrency limit each change the effective unit cost, and states plainly which workloads a speed premium actually pays for.
replicate pricing: when the meter is the clock, not the token
Replicate pricing is the cleanest example of a meter that is literally time, and it is worth starting there because it inverts everything a token card assumes. For private deployments Replicate bills the underlying hardware by the second for all online time — setup, idle and active processing together — and scales to zero when traffic stops, while public models are billed either for active run time or at a flat per-unit rate the model's author sets (a per-image, per-second or per-token figure). The consequence is direct: the cost of a job is dominated by how long the model runs, not by how many tokens it emits, so on a clock-priced meter a slow model is expensive per job and a fast one is cheap even when both would produce identical output.
That is the opposite pole from a per-token speed card. On a token meter the tokens are the quantity and speed is a property of the supply; on a clock meter the seconds are the quantity and speed is the price. Comparing the two means comparing a count with a duration, and the two convert only through an equivalence the provider rarely prints — a tokens-per-second figure you would need before a clock price could ever be expressed as a per-token one. The practical reading is that a clock meter rewards throughput and punishes latency immediately, whereas a token meter is indifferent to latency until you multiply by concurrency.
The two shapes therefore suit different jobs, and it is not a coincidence that they are marketed that way. A scale-to-zero, clock-priced deployment is attractive for occasional heavy work where you control the total online time and would rather not pay for an idle card. A token-priced speed card is attractive for high-request-rate, short-output work where the wall-clock speed decides how much capacity a fixed budget buys. The cluster's taxonomy of billing mechanisms treats "per second" and "per token" as two members of one family; the pricing consequence of that family split is what this section just walked through.
llama api pricing: one set of weights, several speed tiers
Llama api pricing is a useful test case because the weights are open, so the model is not the source of the price difference — the serving is. The same Llama checkpoints ship on many endpoints, and their cards separate on a figure that has nothing to do with the weights: how many tokens per second the endpoint delivers, printed next to the price. A model named for latency rather than for a parameter count is the tell; the naming is a promise about time-to-token, and the card prices that promise.
This is why an open-weight family can carry several prices per million tokens for what a reader thinks is one product. The cheaper row may be a pool that fills a batch; the dearer row is a pool that answers on the interactive path, where the provider holds headroom so that a request is not queued behind another tenant's long job. Holding that headroom is a real cost — idle capacity reserved for latency is capacity that cannot be sold to a batch — and the speed premium is how it is recovered. When you choose the faster endpoint for the same open weights, you are not buying a better model; you are buying a shorter queue. At the bottom of the same family the entry rung is often simply free, and there the limit rather than the rate is the figure to read: a free tier's caps, not its price decide what it can actually serve.
The buyer's move is to stop treating the family as one line item and to ask which tier each call needs. A latency-tolerant bulk job on the same Llama weights belongs on the cheap row and can even be scheduled; an interactive feature belongs on the fast row and pays for the headroom. That split is the practical form of the whole page: the weights are constant, so the only thing left to price is how quickly you insist on receiving them — and, as the next section shows, how quickly you insist on receiving them is partly a question about how many callers are waiting at once.
ai inference cost: what a speed tier moves in the unit rate
Ai inference cost, read from the buyer's side of a card, is a single number: price per token. The speed tier does not change that number — the same tokens cost the same to buy at any speed — so a naive comparison finds no difference to price, and stops. The error is to look only at the numerator. Speed moves the denominator of time: at a fixed request rate, a faster endpoint completes each request sooner, so the same monthly spend clears more sessions, and at a fixed concurrency it removes the queue that would otherwise force you to buy more capacity. On a per-token card, therefore, the observable effect of speed is not a different rate but a different number of requests served per hour of concurrency.
The clean way to hold both facts at once is an effective cost per session, exactly twice: once as tokens per session multiplied by price per token, and once as sessions the provisioned concurrency can serve per hour. The first product is speed-blind and gives the rate card its due; the second is where the speed tier lives, because concurrency multiplied by the inverse of per-request latency is a throughput ceiling. A provider that serves twice the tokens per second raises that ceiling at the same price, which is the same as halving the concurrency you must provision — and provisioned concurrency is the line a platform, not the model, charges for.
Two honesty notes belong here. First, the compute side that actually produces this speed — accelerator hours, batching and utilization — is a different account, and the compute-side cost of inference works it; this page only prices the delivered speed, not the silicon behind it. Second, speed lowers unit cost only while something is waiting on it. With no queue and no interactive caller, faster hardware serves the same tokens for the same money, and the only remaining lever is the batch discount the next sections cover.
Two card shapes make that concrete. A standard-tier card with a discounted asynchronous row prices the same tokens two ways for two timing allowances, so the buyer is choosing a latency budget as much as a rate; and a provisioned deployment that reserves capacity stops selling tokens altogether and starts selling throughput, which is the purest purchase of speed a card can offer. Both are the same idea from opposite ends: where a queue exists, time is the product.
token metering: the meter counts tokens, not seconds
Token metering is the same mechanism whether or not the provider advertises speed, and that is the point. The meter counts input tokens and output tokens; it does not count time-to-first-token, it does not count tokens per second, and it does not charge for the waiting. Streaming a response does not add to the count and hiding it behind a queue does not subtract from it — the tokens are tokens however quickly they arrive. So a speed-priced card and a slow one meter identically, which is exactly why speed has to be expressed somewhere else: in a tier, in a rate limit, or not priced at all.
That blindness has two consequences a buyer should carry into any comparison. The first is that a rate comparison is fair on the meter's own terms only — two providers that count the same tokens at the same rate are equal on the meter, and the difference the buyer feels (a queue, a stalled stream, a request that takes twenty seconds) never appears on the invoice. The second is that the meter hands the provider a lever that is not a price at all: the rate limit. When the tokens are the only thing charged, the way to ration a fast pool is to cap how many requests a key and a team may make per minute, and those caps become the real throughput story — a subject this page returns to with our own numbers below. What a count actually is, tokenizer by tokenizer, is the neighbouring question of how usage is metered and is left there; the pricing point here is narrower, that the meter is blind to speed and therefore cannot be the place where speed is charged for. Even a discounted input class behaves this way: a repeated prefix billed at a published cache-write and cache-read tier is still counted in tokens, and the saving sits in the rate, not in the meter.
embedding pricing: the per-token job with no speed premium
Embedding pricing is the control case for the whole argument, because it is a per-token job with no interactive speed dimension to sell. Embeddings are computed in bulk, nobody waits at a keyboard for a single vector, and the work is naturally batchable and cacheable, so a provider competes on the token rate and on bulk efficiency rather than on latency. There is no queue to jump, no headroom to hold open for the interactive path, and therefore no obvious tier on which to charge a speed premium — the same absence of a queue that made the interactive row expensive to supply is what makes the embedding row plain.
Reading a card with an embedding row next to a chat row is instructive. The chat row carries the promise of time, because its demand is interactive; the embedding row carries only a count, because its demand is throughput. The comparison a buyer should resist is the one that ranks providers by their embedding rate and then extends the conclusion to chat — the two rows are priced under different constraints, and the constraint that dominates chat (waiting callers) is simply absent from embeddings. What an embedding row does share with every other per-token row is the property that its price is volume-driven: more tokens is the only thing that moves the bill, which is the cleanest illustration that a token meter charges for amount and, left to itself, prices nothing about time.
llm token cost: throughput economics inside one chat loop
Llm token cost in an interactive loop is where the two units finally collide, and where the phrase "faster is cheaper" is either true or false depending on the load. Picture a chat feature with a fixed number of active users. Each turn costs tokens at the card's rate — that part is fixed. But each turn also occupies the endpoint for its latency, and a turn that finishes in half the time frees its slot in half the time, so the same provisioned capacity serves more turns per hour. When the binding constraint is concurrency rather than token spend, a faster endpoint lowers the capacity bill at the same token rate, and the effective cost per active user per hour falls even though no per-token figure changed.
The same arithmetic says where the claim fails. If the workload can wait, latency buys nothing and the batch tier's discount buys everything, so the buyer should trade the speed away deliberately. If the workload is long-input, the time is dominated by reading the prompt rather than by emitting the answer, and external tokens-per-second figures — which are usually measured on output — overstate how much of the end-to-end wait the speed premium actually removes. The honest test is a table, not a slogan:
| Workload shape | Binding constraint | Does a speed tier pay? |
|---|---|---|
| Interactive chat, short output | concurrent users | Yes — latency sets capacity, so speed lowers the capacity bill |
| Interactive chat, long input | prompt reading, not decode | Weakly — output speed addresses the smaller half of the wait |
| Batch classification or extraction | throughput, not latency | No — the asynchronous tier's discount dominates |
| Embedding or bulk indexing | amount of text | No — there is no interactive queue to price |
The rule that survives all four rows is that a speed premium is a purchase of capacity headroom, and headroom only has value when callers are waiting. Where none wait, the discount is the better trade. Where the same tokens must be served to many waiting callers, speed is the cheaper way to buy the capacity — which is the mirror image of the mechanism that a route like a cheap open-weight endpoint optimises for: there, the row is priced low on the token rate itself rather than on the speed of each token.
What our own limiter says about throughput pricing
A gateway that prices a platform rather than tokens still has to answer the throughput question, and
the answer lives in its rate limiter — ours is readable in two files, read on 2026-10-08 from
backend/smartgate/core/plan_entitlements.py with rate_limiter.py. The catalog in that file assigns
every plan a REST write rate, an MCP request rate per key, and a separate MCP ceiling per team:
| Catalog plan | REST writes / min | MCP requests / min per key | MCP requests / min per team |
|---|---|---|---|
| FREE | 20 | 30 | 30 |
| PRO | 120 | 300 | 600 |
| TEAMS | 300 | 600 | 3000 |
| ENTERPRISE | 600 | 1200 | 9999 |
Two design choices in that table are the whole argument of this page, measured rather than asserted.
First, the per-key rate and the team ceiling are different numbers, so the throughput a fleet gets
is the smaller of its keys multiplied by the per-key rate and the ceiling itself: a plan whose ceiling
is well under keys-times-per-key stops buying aggregate throughput once the fleet is wide enough, and
the ceiling — not the token price — is what caps how fast a team can spend. Second, the limiter counts
each call in two buckets at once, one per key and one per team, and refuses with a retry hint and a
scope label (mcp_key or mcp_team) so the caller learns which ceiling it hit; when the backing store
is unavailable it fails open, admitting the request rather than blocking traffic on a monitoring
outage. Those are throughput decisions, and none of them touches the price of a token.
The resolver merges the catalog with a team's own feature overrides, caches the result for a minute and falls back to the free plan's limits if the lookup fails, so the effective caps are configuration rather than a hard-coded rate. That is the shape to expect from any throughput-priced platform: the meter that matters is requests per minute per key and per team, and the per-token rate a provider quotes is a different instrument entirely. The published plan figures described below were re-checked against the live pricing page on the same date.
How SmartGate's limits sit next to a per-token speed card
SmartGate does not sell tokens, so it does not sit on the per-token row at all; it prices a platform over the tool traffic that fills an agent's context before a model is asked to complete anything, and it publishes operational limits rather than rates:
| Plan | Monthly token cap | Audit-log retention | Max keys per team |
|---|---|---|---|
| Free | 2M | 7 days | 2 |
| Pro | 20M | 30 days | 10 |
| Teams | 100M | 90 days | 30 |
| Enterprise | 200M+ | 180 days | 9999 |
Read against the argument above, those columns are not a price; they are a set of ceilings, and the per-key request rates published alongside them are the throughput numbers this page has been calling the real constraint. The billing model is the part worth stating plainly, because it is where a platform can avoid behaving like a markup: pay for the platform, and share only when it saves you something — the savings share begins after a measured threshold and the Pro total is capped, so the fee does not grow with every call the way a per-token router's would. Any decision that depends on the caps, the per-key rates, the retention windows or the key counts should be read from the pricing page, which is the authoritative table and this page is not.
How to get started
The first move costs nothing and is the one that keeps the two units from blurring: label the workload before pricing it.
- Write the constraint next to each call. Is the endpoint's cost driven by tokens spent, by concurrency and latency, or by neither? The answer decides whether a speed tier can ever pay.
- Pick the tier the label names. Interactive short output belongs on the fast row; bulk and latency-tolerant work belongs on the batch row, where the discount is the point.
- Price the capacity, not only the tokens. Multiply tokens per session by the rate for the token half, then check the concurrency the endpoint can serve for the same money — that second number is where speed shows up.
- Read the ceilings as part of the price. A per-key request rate and a team ceiling decide throughput as surely as a token rate decides spend; compare them before you commit a fleet.
- If you route tool traffic through a gateway, start free and watch one call end to end, then confirm which ceilings your real load needs on the pricing page.
Frequently Asked Questions
Is Groq pricing cheaper per token because it is faster?
No — the two claims are independent. The tokens-per-second figure is a property of the serving, and the price per million tokens is a separate number on the card. A provider can be fast and mid-priced or slow and cheap. Speed changes how much capacity a budget buys; it does not, by itself, change what a token costs.
When is a faster endpoint actually cheaper?
When callers are waiting and the binding constraint is concurrency rather than token spend. In an interactive loop, a turn that finishes sooner frees its slot sooner, so the same provisioned capacity serves more sessions per hour and the effective cost per active user falls at an unchanged token rate. With no queue, speed saves nothing.
Does the batch discount stack with caching?
At Groq the batch tier halves the standard rate and does not stack with the prompt-caching discount; batch tokens bill at the flat batch rate regardless of cache status, and batch processing does not consume the standard per-model request limits. So a job that can wait should take the discount rather than chase speed.
Why do two endpoints for the same open weights cost different amounts?
Because the weights are open, the price is about serving, not about the model. The cheaper row is often a pool that fills queues, and the dearer row holds headroom so an interactive request is not stuck behind another tenant's job. You are buying a shorter queue and paying for the idle capacity that keeps it short.
What actually limits how fast a team can spend tokens?
Usually the rate limits, not the price. A per-key request rate and a separate team ceiling cap throughput per minute, and when the ceiling sits below keys multiplied by the per-key rate, widening the fleet stops adding throughput. The limit, not the token rate, is the throughput budget.
Limitations
This page treats speed as a pricing axis, and it does so qualitatively on purpose: the provider's own tokens-per-second figures move with the model, the hardware and the load, and a number copied into a comparison is stale the moment the card changes, so the page describes the structure rather than pin a figure. A reader who needs a current rate should read the temperature of the provider's own card on the day they spend against it.
The worked reasoning also assumes a clean workload. Real traffic mixes interactive and batch calls through one key, holds long inputs whose wait is dominated by the prompt rather than the answer, and carries retries and failed requests that a per-token meter may still bill. The "does a speed tier pay" table is a decision aid, not a measurement of any particular system, and a real capacity plan should be built from the buyer's own latency and concurrency numbers.
Finally, the platform limits quoted here are product facts read from the product source and re-checked against the public pricing page on the date above; they move with plan changes, so the pricing page stays authoritative and any budget should be built from it rather than from a copy.
Sources
- Groq Batch API documentation, for the 50 percent asynchronous discount, its non-stacking with prompt-caching and its separate rate limits: https://console.groq.com/docs/batch
- GroqCloud rate-limit documentation, for the per-key and per-team request/token limit structure a token-priced speed card publishes: https://console.groq.com/docs/rate-limits
- GroqCloud model catalog, for the tokens-per-second figures printed beside the per-token rates: https://console.groq.com/docs/models
- Replicate billing documentation, for per-second hardware billing on private deployments, scale-to-zero, and the per-unit rates public models carry: https://replicate.com/docs/topics/billing and https://replicate.com/pricing
- Product behaviour and the plan table: read read-only from the product source at the revision recorded
in this project's
pipeline_results.json, with the plan figures re-verified against the live /pricing page on 2026-10-08. - The rate-limit catalog and the per-key / per-team limiter are our own implementation, read on
2026-10-08 from
backend/smartgate/core/plan_entitlements.pyandmodules/.../rate_limiter.py(origin/main): the four-plan catalog values quoted above, the two-bucket counter, the fail-open behaviour and the sixty-second cache.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher
pinned 0 of 7 sections for this page (0 abstention(s), 7
no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the seven
section keywords, and the remote candidate fallback returned only score-ranked near-misses, which a page
must not dress up as a pin. So every section above but one is written from public sources, each
cited with its own checked date. The exception is "What our own limiter says about throughput pricing":
our own rate-limit catalog and limiter, read read-only on 2026-10-08 from
backend/smartgate/core/plan_entitlements.py with modules/.../rate_limiter.py, stating the four-plan
caps, the two-bucket counter and the fail-open behaviour. The section keyword quoted above each heading
is this project's own measured pool phrase, not a code symbol, and every one carries a measured search
volume above zero. Product claims were read read-only from the product source at the revision the slice
run recorded, with the plan figures re-checked against the live pricing page on 2026-10-08. No code,
batch fingerprints, auction data or internal hosts appear in the text.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|