SmartGateSmartGate

OpenRouter Pricing: Three Components Behind One Model

OpenRouter's own price is not a model rate. The model tokens pass through at each provider's published input and output rates, and the platform's own charge is a fee on the money you load as credits, plus a usage fee once a bring-your-own-key allowance is spent, with a free tier on top.

Short answer: OpenRouter's own price is not a model rate. The model tokens pass through at each provider's published input and output rates, and the platform's own charge is a fee on the money you load as credits, plus a usage fee once a bring-your-own-key allowance is spent, with a free tier on top. So "openrouter pricing" is three components stacked — an upstream unit price, a platform charge on funding or on usage, and whether you bring the key — and the same model can show two effective prices depending on which component is active.

Key takeaways

  • The token rate is the provider's, passed through. OpenRouter describes charging each model's published input and output rates with no markup on the tokens themselves.
  • The platform charge is on funding, not on tokens. The headline figure is a fee applied when you buy prepaid credits, so it scales with the money loaded rather than with tokens consumed.
  • Bring-your-own-key swaps the mechanism to a usage fee that starts after a monthly free request allowance, with the upstream billed to your own provider account instead.
  • The same model can legitimately show two prices — a credits price and a BYOK price — and public reports note the two converging when the routed provider's own card matches the listing.
  • Do this first: write the three components down for each candidate, then price your own workload on the pass-through and add the platform charge as a separate line.

Most of the argument about this topic is an argument about which number is "the price", and the honest answer is that there are three, stacked. A token rate, a fee on funding, and a key-ownership rule are all called pricing, and none of them is wrong — they answer different questions, and a page that leads with one invites a reader to copy it into a spreadsheet where it is silently reinterpreted as another.

A figure quoted per million tokens is a statement about metering usage; a figure quoted as a percentage of a top-up is a statement about funding an account; a free monthly allowance on a key you supply is a statement about routing. That discipline — naming the mechanism before the number — is mapped in full on the cluster's centre, the taxonomy of billing mechanisms; this page applies it to the aggregation and routing market, and stops where a single vendor's own card takes over.

openrouter pricing: three components, not one number

OpenRouter pricing reads like a single rate and is actually three components stacked, and the first useful move is to separate them. The first is the upstream unit price: the input and output rates the model's own provider publishes, applied to the tokens the request uses. OpenRouter describes this as charging each model's published token rates with no markup on the tokens themselves, which is the pass-through claim this page turns on. The second is the platform's own charge, and this is where the money question lives — a fee applied when you add prepaid credits rather than a markup on each token, with a headline percentage on card purchases, a slightly lower one in crypto, a listed business-tier rate, and a small minimum on a top-up. The third is key ownership: with bring-your-own-key, the upstream is billed to your own provider account and the platform charges a usage fee instead, with a free monthly request allowance before that fee begins.

Those three components are why the searches around this topic split into two questions. "How much does a token cost" is about the first component and has the provider's own answer. "What does the platform charge" is about the second and third and has an answer denominated in percent, not in dollars per million tokens. The confusion in public threads — developers asking whether anyone understands the pricing, pointing at an "up to" figure for output — comes from reading one component and expecting the other; a card quoting "up to" a ceiling is quoting the top of a range, not a rate.

The standalone conclusion worth carrying away: an aggregator can be structurally markup-free on tokens and still charge a real fee, because its fee is attached to funding the account rather than to the tokens themselves.

openrouter cost: where the platform fee lands

The cost of a call and the cost of using the platform are two different lines, and conflating them is the most common error in an openrouter cost estimate. The call cost is the pass-through: prompt tokens times the model's input rate plus completion tokens times its output rate, exactly as the model's own provider would bill it. The platform cost is the credit fee, and its defining property is that it scales with how much money you load into the account, not with how many tokens you consume.

The two axes behave differently under load. A per-token markup grows with usage: route twice as many tokens and you pay twice the spread. A fee on prepaid credits does not: you pay it when you top up, and the tokens you then buy are the provider's tokens at the provider's rate. So the effective percentage has a floor set by the top-up fee and a ceiling that depends on how completely you spend each credit purchase — load a large balance to cover a lot of usage and the fee spreads thin; top up in small amounts often and the same percentage recurs. The answer to "how much do a thousand tokens cost" therefore has no single value: the token cost is set by the model, the platform cost by the account funding, and the two meet only through the balance you keep.

Bring-your-own-key changes the second line without touching the first. With your own provider key the upstream calls are billed by the provider on your own account, so the pass-through disappears from the platform's side and is replaced by a usage-based fee that begins after a monthly free request allowance. That is a third shape — not a fee on money loaded, not a markup on tokens, but a charge on the routing itself. A cost model that assumes one of these shapes will be wrong about the other two, which is why the shape comes before the number.

openrouter pricing comparison: the same model at two prices

An openrouter pricing comparison is confusing for a structural reason: the same model can legitimately carry two effective prices at the same moment, and neither is a trick. The first is the credits route — you fund a balance, the platform buys the upstream tokens with it, and your price is the provider's rate plus the amortised top-up fee. The second is the BYOK route — you present your own provider key, the upstream is billed to you directly at the provider's rate, and the platform's charge is a usage fee that only starts after the free allowance. Same request, same model, two contracts, one difference that comes down to who holds the provider key.

A third source of two prices has nothing to do with the platform layer at all. Open-weight models are served by several upstream providers, each with its own rate for the same weights, and a router may select among them. When the router picks the cheaper upstream the pass-through price is genuinely lower than one particular provider's card; when it picks a faster or more available one the price can be higher. A comparison that does not fix which upstream served the request is partly comparing routing policies rather than prices.

The decision rule falls out of the components. The routed price is lower when the platform's own charge is smaller than the spread you would otherwise pay — a free tier, a promotional route, or a router that selected a cheaper upstream. It is necessarily higher when the credit fee, or the post-allowance BYOK usage fee, exceeds the value of the convenience the routing buys, or when the router chose an upstream whose rate sits above the cheapest one available. Under a pure pass-through with no fee at all the two prices converge — the tell that the first component is a genuine pass-through and any difference lives in the second.

llm price comparison: pass-through versus marked-up

A llm price comparison gets easier once you ask one question about each candidate: is this number the provider's own price, or the provider's price plus someone's spread? A pass-through price equals the rate the model's provider publishes for the same model identifier. A marked-up price is that rate plus a spread the reseller keeps, and the spread is often invisible because the reseller publishes one blended figure rather than the components that produced it.

Three cheap tests. First, compare the listing against the provider's own card for the same model id: equal means pass-through, higher means a markup, and a note that the token is served by a partner upstream means the comparison moved to a different provider. Second, look for a revenue line that is not the token rate — a fee on credits, a subscription seat, a per-request charge. A service can pass tokens through at cost and still earn on one of those, which is the shape this page's subject takes. Third, ask whether the platform lets you supply your own key; if it does, it has told you which side of the trade it is on for the upstream cost, because a BYOK path hands that billing back to you.

The reason this matters is that "pass-through" is a claim about one component, not about the total. A service that passes tokens through at the provider's exact rate and charges a fee on your credits is not automatically cheaper or more expensive than a service that marks tokens up; it is a different mechanism, and the comparison has to be run on the same mechanism or it ranks nothing. The provider-side cards these tests compare against are documented on this cluster for OpenAI's standard tier and DeepSeek's per-token rates; the pass-through claim is verifiable only against those, model id by model id.

llm pricing comparison: reading a routed rate

An llm pricing comparison across routed services has one extra variable a direct comparison does not: the route. When a request can reach a model through several upstreams, the rate on the invoice is the rate of whichever upstream served it, so a fair comparison holds the route constant before it compares anything else. Fix the model identifier, fix the route, fix the token counts, and only then line up the platform charge — a comparison that lets any of the four float is measuring the float.

Cached input is the second variable a routed rate exposes, and the one most often averaged away. A provider that has already processed a prefix can re-read it at a fraction of the input rate on a later call, and a pass-through router carries that discounted cached rate through. Two otherwise identical workloads can therefore differ in effective input price purely by how much of their context is stable. The same detail also behaves differently from provider to provider: the cache tiers on Claude's pricing are an explicit, priced write with a lifetime, while some providers apply an automatic prefix discount, so a cached rate on a card and the cached rate on your own invoice are not automatically the same claim.

A third check is availability, which is not a price but does change which price you pay: a router that offers a fallback when the preferred upstream is unavailable will sometimes serve a request at a different rate than expected, and a ranking on the cheapest listed rate that ignores the fallback path is ranking a route that may not be used. The honest deliverable is a ranked cost for a named workload on a named route, with the fallback noted — the frame on Groq's per-token rates as much as on any routed listing.

model pricing comparison: why identical model IDs differ

A model pricing comparison looks objective — same model, two cards, one number each — and is usually comparing more than the model. The same identifier can carry different rates because of five things around it, each a legitimate reason for a difference rather than an error.

  1. Which upstream serves it. An open-weight model is offered by more than one provider, at more than one rate, for the same weights, so the identifier alone does not determine the price.
  2. The context-length tier. Above a stated prompt length some providers reprice whole token classes, so the headline rate applies only inside the smaller window.
  3. Cached versus uncached input. A stable prefix may be billed at a fraction of the input rate, moving the blended cost without moving the number printed on the card.
  4. The batch or interactive shape. A queued request can be billed below the interactive rate because the provider schedules it on hardware it would otherwise idle, trading latency for money.
  5. The platform layer. If the request is routed, the figure on the invoice may carry a credit fee, a usage fee, or a spread that is not a property of the model at all.

The fifth item is the one to be most careful with, because the first four are visible on a card and the fifth often is not. The useful habit is to write the model identifier and the route next to each number and compare only numbers that share both; a rate with no route beside it is an unfinished comparison. Where the route is a genuinely free tier rather than a paid one, the constraint moves from price to limits — the request rate, the daily allowance, the model list — which is its own question, covered on free-tier LLM APIs. Price is the wrong axis to compare a free tier on at all; its limits are the whole of its shape.

ai pricing comparison: what to line up before the figure

An ai pricing comparison is only a comparison when both sides charge by the same mechanism, and for routed services that means lining up three components before reading a single rate. Write them down in a fixed order.

Component What it is How it scales Where it hides
Upstream unit price the provider's input and output rate for the model with tokens consumed which upstream served the call
Platform charge a fee on credits loaded, a usage fee, or a subscription with money loaded, or with routed requests the funding step, not the invoice line
Key ownership whether you bring your own provider key with the post-allowance usage fee the free allowance that precedes it

Two rules follow. Compare components one at a time: match upstream rates against upstream rates and platform charges against platform charges, rather than comparing a per-token figure from one service against a percentage fee from another, because those two numbers do not share a scale. And state the workload before the ranking: a light, bursty user and a heavy, steady one face different effective rates from the very same fee schedule, because a credit fee is amortised over money loaded and a usage fee over requests, so one schedule rewards opposite traffic shapes. A comparison that names its workload can be trusted; one that names a single winner across all workloads cannot, because the winner is a property of the workload as much as of the vendor.

The subscription shape some tools use is a fourth mechanism beside these, and it is the one the public questions reach for when they ask what the best AI subscription for the price is, or whether a percentage rule of thumb applies to AI spend. Those are questions about a flat fee against a usage forecast, not about token rates.

What our own price table looks like

Our own budget table is a useful contrast because it is the opposite shape from a routing market. Read on 2026-10-08 from origin/main, the table is a local JSON file, backend/smartgate/modules/budget_guard/model_prices.json, holding five models at that revision. Each model carries exactly four values: an input cost per token, an output cost per token, a maximum input length and a maximum output length. There is no column for a platform fee, no column for a credit spread, no flag for whether a key was supplied by the caller, and no per-route variation. The unit is one token, not a price per thousand, so there is no implicit division hidden between the stored rate and the arithmetic that uses it.

The arithmetic lives in the same folder, in modules/budget_guard/algorithm.py. Counting chooses a tokenizer by model family — one encoding for the GPT-4o and DeepSeek families and another otherwise, with a fallback — and costing is the provider-neutral formula: the input rate times the prompt count plus the output rate times the completion count, with an optional caller-supplied override for a custom price. The module records that its counting core is extracted from an open-source client with the provider-specific pricing removed in favour of one unified table.

The honest reading of that shape is narrow: a table with one rate per model and no fee column is the implementation form of a price that does not depend on the route the request took, because no route is recorded to charge differently. We do not operate a marketplace over model tokens and do not take a spread on them; the table simply has nowhere to put such a spread. What the platform does price is the tool traffic it handles before a model is asked to complete anything, described below. The figures carry the date they were read, so a budget built on them should re-read the file rather than trust this page.

How SmartGate sits next to a routing market

SmartGate is not a routing market and not a reseller of model tokens, so it does not fit the component stack above. It is a gateway over the tool traffic that fills an agent's context before a model is asked to complete anything: fetching, searching, compressing a long context, de-duplicating repeated material, checking a budget, storing team memory, and orchestrating the steps. Those calls are what the platform's plans meter, and the billing model is deliberately not a per-token markup — the platform is paid for, and a share is taken only once the platform has measurably saved you something, with the Pro tier capped so the fee does not grow with every call. The operational limits rather than rates sit on the plan table: monthly token caps rise across the tiers, the per-key request rate rises with them, and audit-log retention and per-team key counts follow. Those numbers move with plan changes, so the pricing page is authoritative and this paragraph is not.

Read against the three components, the placement is clean. There is no upstream model unit price, because the platform does not resell model tokens. There is a platform charge, but it is a savings-share rather than a fee on funding or on requests, so it does not scale either way. There is no key-ownership component, because the platform is not competing for your provider credentials. If your problem is the price of a routed model, the components above are the ones to compare; if your problem is the tool traffic that surrounds the call, that is the layer this platform prices.

How to get started

Three steps, in the order that saves the most confusion.

  1. Name the three components for the service in front of you: upstream unit price, platform charge, key ownership. Write "unknown" where a service does not say, rather than quietly assuming markup-free.
  2. Price your own workload on the pass-through alone — prompt count times the input rate plus completion count times the output rate — and then add the platform charge as a separate line. Never let the two blend into one number, because that is the step where the comparison stops being like-for-like.
  3. Run one real call end to end and read the actual charge, then confirm the tier your steady traffic needs. To watch a call pass through a gateway, start free, and read the pricing page for the plan limits.

The exercise is worth an hour: it replaces a remembered headline percentage with a number computed from your own token counts.

Frequently Asked Questions

Is OpenRouter's price the same as the provider's price?

Not as a total. The token rate is the provider's published input and output rate, passed through with no markup on the tokens themselves; the platform's own charge is separate, either a fee on the credits you load or a usage fee once a bring-your-own-key allowance is spent. So the token rate can match the provider's card exactly while the amount you pay is the provider's rate plus the platform charge.

Why does the same model cost two different amounts?

Because there are two funding routes and sometimes more than one upstream. On the credits route you fund a balance and the platform buys the tokens with it, so you pay the provider rate plus the amortised top-up fee. On the bring-your-own-key route the upstream is billed to your own account and the platform charges a usage fee after a free allowance. And an open-weight model may be served by several providers at different rates, so the route the router selected also moves the number.

Does the platform charge a markup on tokens?

It describes the tokens as passed through at the model's published rates with no markup, and its own charge as a fee applied when you buy credits, plus a usage fee on the bring-your-own-key path. That is a different mechanism from a per-token markup: a markup scales with tokens consumed, while a fee on funding scales with the money you load.

When is a routed price actually cheaper?

When the platform's own charge is smaller than the spread you would otherwise pay — a free tier, a promotional route, or a router that selected a cheaper upstream than you would have picked. Under a pure pass-through with no fee the routed price and the provider's price converge. It is more expensive once the credit fee, or the post-allowance usage fee, exceeds the value of the convenience.

How much will a thousand tokens cost?

There is no single answer, and that is the point. The model decides the token cost, through its provider's input and output rates; the platform charge is attached to how you fund the account rather than to the token count, so it does not divide into a per-token figure. Compute the pass-through for your own token mix and add the platform charge as a separate line.

Limitations

This page describes the structure of an aggregator's price and deliberately does not reproduce a moving rate as a quote. The fee percentages and free-allowance terms it names are the figures reported at the checked date; they keep their own source and should be re-verified on the provider's own pricing page before any spend decision, because a percentage that changes turns a copied number into a stale one. The page also does not compute a bill, does not cover cloud-hosted resale of a vendor's models, and does not read any single vendor's card — those are separate sections of the same cluster, linked where the argument needs them, and each owns its own mechanism.

The reading of our own table is scoped to what the file holds: five models at one revision, four values each, read on the date above. It is evidence about the shape of our implementation, not a claim about any other service, and it says nothing about a rate outside those five rows. Where a reader needs an actual figure for a specific model, the honest path is the provider's own table on the day of the decision.

Sources

  • OpenRouter's own pricing and models pages — https://openrouter.ai/pricing and https://openrouter.ai/models — for the pass-through token rates, the credit fee, the bring-your-own-key terms and the free tier described above, checked 2026-10-08.
  • The AI Overview and the third-party guides the captured SERP for the head phrase surfaced (openrouter.ai, truefoundry.com, amnic.com), which state the credit-fee percentages and the BYOK allowance quoted here, read 2026-10-08.
  • The public developer discussion on the pricing of routed models (the Hacker News thread in this project's serp/openrouter-pricing.json) for the reported convergence between a routed rate and a provider's own card for some models, read 2026-10-08.
  • Our own implementation: backend/smartgate/modules/budget_guard/model_prices.json with modules/budget_guard/algorithm.py, read read-only from origin/main on 2026-10-08.
  • Demand context: the section phrases come from this project's own measured pool (search_volume.json / research_brief.md, DataForSEO Google Ads, measured 2026-10-08), not a third-party keyword tool.

Method note

This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of the 7 planned sections for this page: rule A found no unique symbol in the scanned repository for any section keyword, and the remote candidate fallback returned only score-ranked near-misses. Every section above is therefore written from public sources — the provider's own pricing page and the guides the SERP surfaced, cited with a checked date — because the house rule for an unpinned section is sourced, never invented. The one exception is "What our own price table looks like": our own price table and counter, read read-only on 2026-10-08 from backend/smartgate/modules/budget_guard/model_prices.json with modules/budget_guard/algorithm.py, stating four values per model, the per-token unit and the per-family tokenizer choice. The section keyword quoted above each heading is this project's own measured pool phrase, not a code symbol, and every one of the seven carries a measured search volume above zero. No code, batch fingerprints, auction data or internal hosts appear in the text.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)