OpenAI API Pricing: Generations, Entries and Units
OpenAI API pricing is a single vendor's list, but no single number describes it. The same model carries a different effective rate at each entry point you call it through — the synchronous chat surface, the responses surface, the batch queue, and cached input — and a new model generation moves the whole price band rather than one row.
Short answer: OpenAI API pricing is a single vendor's list, but no single number describes it. The same model carries a different effective rate at each entry point you call it through — the synchronous chat surface, the responses surface, the batch queue, and cached input — and a new model generation moves the whole price band rather than one row. The per-million figure printed on the card and the per-token figure stored in a code table are the same rate written in two units.
Key takeaways
- The card is one vendor, not one price. OpenAI's list is a grid of models by generation, and each model's line splits again by entry point, so "the OpenAI API price" is a family of rates.
- A generation is a band, not a row. When a new family ships, the previous family stays on the card at its own rate band; anchoring a budget on one row means budgeting against one band.
- The entry point decides the rate for the same model. Chat, responses, batch and cached input are four ways to buy the same tokens, and the batch and cached rows are cheaper because you trade latency or cache behaviour for money.
- Per-million and per-token are the same number. Multiply or divide by one million; a card in one unit is never cheaper than a table in the other.
- Do this first: name the entry point and the generation before you compare two figures, then convert both to a single unit — that pair, not the headline, is what a workload is billed at.
The question "what does the OpenAI API cost" has a stable structure and shifting numbers. The structure is what a reader can carry between releases: a list that is organised by model generation, with each model priced at more than one entry point, and a unit convention that lets the same rate appear as a per-million figure on the marketing card and a per-token figure in a library. The numbers move on their own schedule, which is why this page teaches the structure and dates every figure it does quote.
The most common misreading is to treat one row as "the price". A reader who copies the flagship input rate into a spreadsheet and stops has budgeted against one generation at one entry point, and a real workload rarely lands on exactly that cell. The workload calls a specific surface, may accept a queue, and may reuse a prefix across calls; each of those choices selects a different rate on the same card.
This page is one page in a cluster about what AI costs, and it owns the single-vendor view: how one list behaves across generations and entry points. Reading the general method for any card, doing the arithmetic of one bill, classifying the billing mechanisms across providers, counting the usage those rates apply to, and costing your own hardware are five other pages, each linked where a section below actually needs it, and none of them is restated here.
openai api pricing: one card, several entry points
OpenAI's list is a grid, and the first axis is the model. The second axis is the entry point — the surface you call and the timing or state you accept in exchange for the rate:
- The synchronous chat surface. A request that returns in a single call. This is the default rate the card leads with, and the one most comparisons quote.
- The responses surface. A stateful call that keeps the conversation server-side and bills the same token classes as the chat surface, so it changes the shape of the request more than the price of a token.
- The batch surface. Work submitted into a queue for the provider to schedule later, billed below the interactive rate because the provider trades your latency for its own utilisation.
- Cached input. A prefix the provider has already processed is re-read at a fraction of the input rate, so on a repeated-context workload the input line shrinks without any change to the output line.
- Non-token surfaces. Image, audio, transcription and fine-tuning rows are priced per image, per second, per character or per training step — units that do not fold into a token comparison without the provider's own stated equivalence.
The point of separating these is that they are not competing brands of one product; they are the same provider selling the same tokens under different contracts. A model row and an entry-point row multiply, so the price you pay is the model you chose crossed with the way you called it. That is a narrower statement than the general mechanism taxonomy, which lives at the cross-provider view of billing mechanisms; here the axis is one vendor's own list, and the useful first question is which cell of the grid a workload lands on.
Two of the entry points deserve a plain warning. The batch row is a different latency contract, not a discount you can take on interactive traffic — a chat feature that answers a user cannot use it. And the cached-input row is conditional on a cache hit, so it is a rate you can reach only by structuring a workload around a stable prefix; a provider that caches automatically decides when that happens. Read together, the entry-point axis explains why two teams on the same model can see materially different bills.
openai api cost: the generational phase between two bands
An openai api cost estimate is usually an estimate for one generation, and that is where it ages fastest. When a new family ships, the provider does not overwrite the old rows; it lists the new family beside them, and the two sit at different bands. The gap between those bands is a phase offset between generations, and three consequences follow.
First, the headline moves without the previous model changing. A team that ships a feature on the older generation keeps its rate while the front of the card reprices around it, so a competitor who quotes "the current OpenAI price" may be quoting a band that team never pays. Second, the new band is not automatically cheaper across every entry point; a generation can lead on the interactive rate and still be beaten on the batch rate by the model it replaces, because the batch multiplier is applied per generation. Third, because both bands are live at once, a comparison that mixes a new model's interactive rate with an old model's batch rate is comparing two different contracts and calling the result a generation gap.
The practical rule is to date a band, not a model. Write the generation, the entry point and the date you read the number, and re-read it when the front of the card changes — which is often. Where one vendor prices the same underlying model through more than one service, the bands diverge again: a cloud reseller can offer a provisioned rate which reprices by reserved capacity rather than by tokens (Azure's provisioned and regional bands), and a rival vendor may deliberately list a whole generation under this band to win price-sensitive traffic (a provider priced far below this band). Neither changes the reading method; both are the same phase between bands, drawn on a different axis.
The mistake the phase framing prevents is treating a price as a property of the model's name. It is a property of the name plus the generation snapshot plus the entry point. A static table cannot hold that; only a dated read can.
gpt cost: which entry point a request lands on
"gpt cost" as a single figure hides the entry point, and the entry point is the variable a reader controls. Take one model generation and hold it fixed; the same tokens are then available at up to four effective input rates and two output rates, because batch and caching act on different halves of the count:
| Entry point | Effect on input | Effect on output | What it costs you |
|---|---|---|---|
| Synchronous chat | full input rate | full output rate | nothing — this is the baseline |
| Responses (stateful) | full input rate | full output rate | request shape, not price |
| Batch | discounted input | discounted output | latency you must be able to accept |
| Cached input | reduced input on a hit | output unchanged | a stable, reusable prefix |
The asymmetry in that table is the thing to notice. Caching moves the input line and cannot touch the output line, because a token the model has not generated yet cannot be cached; batching moves both lines and therefore changes the whole call. So a retrieval-heavy workload that reuses a long prefix can shrink its bill through the cached row alone, while a generation-heavy workload has less room to move unless it can also accept a queue. The current documentation describes the cached discount as a large reduction of the input rate and the batch tier as roughly half the standard rate; the exact multipliers move between generations, so treat them as the shape to look for on the current card rather than as fixed constants.
What this section does not do is turn those rates into a bill. The multiplication of a rate by a counted mix is its own page (the arithmetic of a single bill), and this page stays on the card: which rate applies, for the model you chose, at the entry point you called. Where a vendor instead sells the tokens through a subscription or an explicit cache-write you pay for up front, the entry-point question changes character, which is the case on Claude's cache-write and seat-based pricing and is deliberately not the subject here.
gpt-4o pricing: one generation's row on the card
gpt-4o pricing is a good worked example of a single generation's line, and the useful thing about it now is not the number but the position. It names one family, and on the card that family has an input rate, an output rate, a batch row and a cached-input row, all sharing one set of context limits. Reading it as a generation rather than as "the price" is what keeps a budget honest once a newer family is listed above it.
Two properties of that kind of row are worth carrying between releases. The first is that the row's output rate is several times its input rate, which is the structural asymmetry of every current token card and not a quirk of this model; a workload that reads a lot and writes little and one that writes pages from a short prompt feel the same row very differently. The second is that the row's limits are part of the price: a context window of a given size and a maximum output of a given size decide which requests are even admitted, so two rows with identical rates can be different products if their windows differ.
This page does not reproduce OpenAI's current per-model figures for this family, because those move and a copied number is stale the day it is printed. The one place a concrete gpt-4o figure appears below is our own stored table, dated, because it is a number we can actually point at. The reading rule is the deliverable here: a generation's row is a band with limits attached, and a budget should name the generation it was built on.
gpt-4.1 pricing: the next band shifts the grid
gpt-4.1 pricing is the next band on the same card, and its value in a reading is that it shows the grid shifting rather than a single row changing. When a successor family arrives, three things move at once: the rate band of the new rows, the position of the old rows relative to the new ones, and sometimes the set of entry points the new family offers first. A reader who only compares the two input rates has compared one cell of two bands.
The honest way to place a new band is a small table with one row per generation and one column per entry point, every figure carrying the date it was read. Ranked that way, the answer is rarely "the new generation wins everywhere". A new family may lead on the interactive rate and lag on the batch rate of its predecessor; it may lead on the flagship row and lose on the small row that most classification traffic actually uses. The same-generation comparison drawn on serving hardware rather than on the model name makes the point from another angle: a provider running a model on specialised inference hardware can list a lower per-token rate for a comparable class (specialised inference hardware that repriced the same class), which is the grid shifting under a fixed model name rather than a family release.
So gpt-4.1 pricing is read the way any new band is read: confirm the unit, identify which entry points the new rows carry, compare only within an entry point, and date every cell. A page that lists one number per model, however current, has flattened both axes of the grid.
gpt-4o cost: per-million and per-token on the same card
A gpt-4o cost figure appears in two places that look different and are the same: a provider's card quotes a price per million tokens, and a library or a local table stores a price per token. The conversion is a single factor of one million, and nothing else about the two numbers disagrees. A per-token rate of 0.0000025 is $2.50 per million tokens; a per-1K rate of $0.0025 is the same again. A comparison that lines a per-million card against a per-token table and finds one "cheaper" has compared a unit, not a price.
The reason both forms persist is that each is convenient where it is used. Per-million is readable for a human scanning a marketing page and matches how the monthly bill is described. Per-token is the form a program multiplies by a count, because it removes an implicit thousand from every calculation — the moment a rate is wrong, it is wrong as a number rather than as a scale. The conversion is worth doing once per card and writing down, because it is the step most often skipped when a figure is copied from one surface to another.
Two rows on the same card can also be compared only after the unit matches, and the comparison is then about more than the headline. The input rate and the output rate convert independently, and so do the cached and batch rows, so a card is a small set of conversions rather than one. A figure that is "per million tokens, cached input" is a different conversion from "per million tokens, input", and treating the two as one number is the same error as treating the chat and batch rows as one. Where a layer resells the same models, it may quote a single blended per-million rate that already mixes input, cached input and output for you (a marketplace that resells the same models at blended rates); that rate is a summary of the reseller's mix, and it converts to per-token just as cleanly, which is exactly why it looks comparable when it is not.
input token cost: our own read of one production table
Every rule above ends in the same place — a rate multiplied by a count — so it is worth showing a
table we can actually point at, and the two OpenAI rows in it. Our price table is a local JSON file,
read here from origin/main on 2026-10-08, and it holds four values per model: input cost per
token, output cost per token, a maximum input and a maximum output. The two OpenAI rows at that
revision read:
| Model row | Input (USD / token) | Output (USD / token) | Max input | Max output |
|---|---|---|---|---|
gpt-4o |
0.0000025 | 0.00001 | 128,000 | 4,096 |
gpt-4o-mini |
0.00000015 | 0.0000006 | 128,000 | 16,384 |
Three properties of those rows are the ones to copy, whatever the numbers are on the day you read
them. First, the unit is one token, never per 1K: the input row for gpt-4o at 0.0000025 per
token is the same rate as $2.50 per million tokens, and storing it per token means there is no
hidden division between the table and the arithmetic that consumes it. Second, the two rows in one
family share an input limit of 128,000 tokens but not an output limit — gpt-4o-mini allows 16,384
output tokens against gpt-4o's 4,096 — which is the limits-are-part-of-the-price point from the
gpt-4o section made concrete on our own numbers. Third, the counter chooses its tokenizer per model
family: a model name containing gpt-4o (or deepseek) is counted with the o200k_base encoding,
and everything else falls back through encoding_for_model to cl100k_base, so the same string is
a different number of tokens on the two rows of a comparison unless the counts come from the right
encoding.
The table is a snapshot, and that is the property to carry rather than a defect to hide. A local, diffable file means a rate change is a reviewable commit rather than a deploy, and it equally means the table is only as current as its last edit. Our own rule applies to us: the figures above carry the date they were read, and any budget built on them should re-read the file rather than trust this page.
How SmartGate's platform sits next to a per-token card
A per-token card prices a completion; SmartGate does not sell tokens, so it does not occupy any row of that grid. It is priced as a platform over the tool traffic that fills an agent's context before a model is ever asked to complete it, and it publishes operational limits rather than rates:
| Plan | Monthly token cap | MCP requests/min per key | Audit-log retention | Max keys per team |
|---|---|---|---|---|
| Free | 2M | 120 | 7 days | 2 |
| Pro | 20M | 300 | 30 days | 10 |
| Teams | 100M | 600 | 90 days | 30 |
| Enterprise | 200M+ | 1200 | 180 days | 9999 |
Read against the card, those columns are not rates at all — they are the ceilings and retention windows that decide what a workload is allowed to do. The billing model is the part worth stating plainly, because it is where a platform avoids behaving like a markup: pay for the platform, and share only when it saves you something — the share begins after a measured savings threshold and the Pro total is capped, so the fee does not grow with every tool call the way a per-token markup would. The caps, per-key rates, retention and key counts change by tier, so the pricing page is the authoritative table and this page is not.
How to get started
- Name the generation before the number. Write which family each figure belongs to, and mark the older band you are still paying on.
- Name the entry point. Interactive, batch or cached input is a different contract; a rate without its entry point is not comparable to another rate.
- Convert both sides to one unit. Move per-million and per-1K to per-token, or the reverse, and leave per-image and per-request rows out until the provider states an equivalence.
- Read the limits as part of the price. A context window and a maximum output decide which requests a row admits, so two same-rate rows can be different products.
- If you route tool calls through a gateway, start free and watch one call end to end, then confirm the tier your real volume needs against the current plan limits on the pricing page.
Frequently Asked Questions
Is there one price for the OpenAI API?
No. The list is a grid of model generations, and each model is priced at more than one entry point — the interactive surface, a batch queue and cached input — so a single figure describes one cell of that grid, not the API as a whole.
Why do two rows for the same model end up in a comparison?
Because the entry points multiply with the model. A batch row and a cached-input row for one model are different contracts, and lining one against another model's interactive rate compares two different things and calls the result a price difference.
Why do a card and a code table seem to disagree?
Almost always a unit, not a price. A rate per million tokens and the same rate per token are one number scaled by a million. Convert both to one unit before comparing, and a remaining difference is a genuine rate difference.
Does a new generation make the old one cheaper?
Not by itself. A new family is listed beside the old rows at its own band, and the old rows keep their rate until the provider changes them. A generation gap is a difference between two live bands, so a budget should name the band it was built on.
Does caching help every workload equally?
No. Cached input reduces the input line on a cache hit and cannot touch the output line, so a workload that reuses a long stable prefix benefits most. A generation-heavy call that writes more than it reads has little to gain from the cached row.
Limitations
This page teaches how one provider's card behaves across generations and entry points; it is not a price list, and it deliberately avoids copying any OpenAI figure it could not date. The one concrete OpenAI rate quoted above is our own stored table's number, read on the date given, and it may no longer match the live card — the provider's own pricing page always wins over any summary, including this one. Batch multipliers, cached-input discounts and context tiers move between generations, so the shapes described here should be re-read against the current card before they enter a budget.
The page also stops at the boundary of the single vendor. It does not compare OpenAI against another provider's rates, does not compute the cost of one request from a counted mix, does not classify the billing mechanisms that all providers share, and does not price self-hosted hardware — those are separate pages in this cluster, linked through the sections above, each owning its own angle. A comparison that needs another vendor's numbers should take this reading method to that vendor's card and re-check it on the date it plans to spend.
Finally, the plan limits quoted here are product facts read from the product source and re-checked against the public pricing page on the date above; they move with plan changes, so the pricing page remains authoritative and any budget should be built from it rather than from a copy.
Sources
- OpenAI API pricing documentation, for the model-and-entry-point structure of the list, the batch tier and the cached-input rate: https://openai.com/api/pricing/ and https://platform.openai.com/docs/guides/batch , checked 2026-10-08.
- OpenAI prompt-caching documentation, for the cached-input behaviour and its dependence on a cache hit: https://platform.openai.com/docs/guides/prompt-caching , checked 2026-10-08.
- OpenAI Batch documentation, for the asynchronous tier as a latency-for-rate trade: https://platform.openai.com/docs/guides/batch , checked 2026-10-08.
- Product behaviour and the plan table: read read-only from the product source at the revision
recorded in this project's
pipeline_results.json, with the plan figures re-verified against the live /pricing page on 2026-10-08. - The OpenAI rows and the counting note above are our own implementation, read on 2026-10-08 from
backend/smartgate/modules/budget_guard/model_prices.jsonandmodules/budget_guard/algorithm.py(origin/main): the per-token input and output rates, the context limits, the per-token unit and the per-family tokenizer choice.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice
matcher pinned 0 of the 7 planned sections for this page: rule A found no unique symbol in the
scanned repository for any section keyword, and the remote candidate fallback returned only
score-ranked near-misses, which a page must not dress up as a pin. Every section above is therefore
written from public sources — the provider's own pricing documentation, cited with a checked date —
because the house rule for an unpinned section is sourced, never invented. The one exception is
"input token cost: our own read of one production table": our own price table and counter, read
read-only on 2026-10-08 from backend/smartgate/modules/budget_guard/model_prices.json with
modules/budget_guard/algorithm.py, stating the per-token input and output rates, the context
limits and the per-family tokenizer choice. The section keyword quoted above each heading is this
project's own measured pool phrase, not a code symbol, and every one of the seven carries a measured
search volume above zero. No code, batch fingerprints, auction data or internal hosts appear in the
text.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|
Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before publication. 0 of 7 sections pinned, 0 abstentions, 7 no-slice sections (all seven written from the provider's published documentation and our own stored table, cited above).