SmartGateSmartGate

DeepSeek Pricing: The Shape of a Low-Price Tier

DeepSeek pricing is the clearest example of a low tier built from a shape rather than a discount: two input prices for a cache hit and a cache miss, an off-peak window that pays you to wait, and a family of models at different rates. The headline figure is low, but it prices a specific workload — short prompts, high cache reuse, work that can wait, and mostly input.

Short answer: DeepSeek pricing is the clearest example of a low tier built from a shape rather than a discount: two input prices for a cache hit and a cache miss, an off-peak window that pays you to wait, and a family of models at different rates. The headline figure is low, but it prices a specific workload — short prompts, high cache reuse, work that can wait, and mostly input. Move off that shape and the same card stops being cheap.

Key takeaways

  • The low figure is an input rate, and input is the cheap direction. On almost every hosted API output is priced above input, so a low card is low because it is weighted toward the direction that is cheap to serve.
  • A cache hit and a cache miss are two prices for the same input token. The effective input rate depends on how much of your context repeats, which is a property of your workload, not of the card.
  • Off-peak pricing is a scheduling trade. A lower rate in a window outside peak hours is real money for work that can wait and worth nothing to a request that must answer now.
  • A model family is a ladder, not a rate. The small member and the reasoning member are different products, and picking the rung a task actually needs is the largest cost decision on a low tier.
  • "Cheapest llm api" is a question about your loop, not about a headline number. The lowest card can carry the highest effective cost for the wrong workload shape.
  • Do this first: write down which of the four dials your workload moves, then compare only cards that respond to the same dial in the same way.

Most pages about cheap models publish a number and stop. The number is real, but on a low-price tier the number is the least transferable part of the card, because four pricing dials move underneath it. The first is the direction of the token: input and output are billed separately on almost every hosted API, and a low tier is usually low because it is dominated by the cheap direction. The second is the cache: a prompt prefix the provider has already processed can be re-read at a second, lower input rate. The third is time: an off-peak window sells the same tokens below the interactive rate because it lets the provider schedule the work. The fourth is the family: a vendor rarely sells one model at one rate, but a ladder of them.

In plain terms for a reader who is not going to read a rate card line by line: the price you are quoted is a contract for a particular shape of work, and the cheapest-looking card is cheap only while your work keeps that shape. This page takes the shape apart and shows where it breaks.

The angle here is deliberately narrow. It is not the list of free or cheap providers — a sibling page owns that — and it is not the arithmetic of a bill. It is the shape of the low tier: what makes a low rate low, which four dials a buyer can actually move, and the workload shapes in which the same low card stops being cheap. This page is a member of a cluster about what AI costs; the centre maps the taxonomy of how providers charge, and this page stays on the low tier rather than restating that catalogue.

deepseek pricing: the shape behind a low headline rate

DeepSeek pricing reads best as a shape with four dials rather than as a discount off another vendor's card. The dials are the same four that any low-price tier in this market turns, and the reason to name them is that a buyer can move three of them and cannot move the fourth.

The first dial is direction. Hosted APIs bill input and output separately, and the low number on a low tier is almost always the input rate: reading a prompt is a parallel operation, generating is sequential, and the published asymmetry is the provider passing that difference through. A card built around a low input rate is therefore a card built for input-heavy work — classification, extraction, retrieval-augmented questions answered in a sentence — and the same card is unremarkable the moment a workload writes long output.

The second dial is the cache. A prefix the provider has already processed can be re-read at a second, lower input rate, so the effective input price of a chat or agent loop depends less on the printed rate than on how much of the context repeats. Cached input is the mechanism that makes a low tier genuinely low for a stable prompt and leaves it at the miss rate for a prompt that never repeats.

The third dial is time. An off-peak or batch window sells the same tokens below the interactive rate because it lets the provider schedule the work on hardware it would otherwise idle — a trade of latency for money that only helps work which can wait.

The fourth dial is the family. A vendor rarely sells one model at one rate; it sells a ladder, from a small, cheap member tuned for short outputs to one or more larger, dearer members that do more per token. "The price" is really a choice of which rung a task needs.

This page does not reproduce DeepSeek's current rates, because a rate that was not re-checked on the day of reading is exactly the kind of number a price page should refuse to copy; the provider's own card is authoritative, and the sections below describe the shape it expresses. The one place this page does quote exact figures is our own price table, read read-only from the product source on 2026-10-08, further down.

cheapest llm api: what the low rate actually prices

The phrase cheapest llm api usually means the lowest published per-token figure, and the first thing to notice is how little that figure commits to. A per-token rate prices exactly the tokens the provider counts on a call it completes. It does not price the tokens an application spends on retries, on tool calls that refill the context before the model is asked anything, on the re-sent prefix after a cache miss, or on a request the provider bills even though the answer was discarded. On a feature-rich tier those are a rounding error against a large completion; on a low tier they are a larger share of the real spend, because the loop, not the single call, is what scales.

That is why the lowest card can carry the highest effective cost for the wrong workload. A workflow that reads a short prompt and writes pages is dominated by the output direction, where a low tier does not compete; a workflow that reads long, stable documents and writes a label is dominated by the input direction, where it does. There is a second ambiguity in the phrase, and it lives above the provider's own rate. A gateway, a router or a marketplace can quote a figure that is not the provider's rate at all — it is the provider's rate plus or minus the middle layer's own margin, and the middle layer may resell some routes below list and others above it (a router that resells at a spread). So the honest version of the question is not who prints the smallest number but which middle layers a given number includes. The mechanism-by-mechanism version of that argument is the centre page's own job, and the point to carry here is narrow: compare only figures that price the same tokens at the same layer.

cheapest ai api: the cache tier is two prices, not one

On a low tier the single most consequential hidden variable is the cache, because it turns one printed input rate into two prices. A provider that keeps a processed prefix can re-read it on a later call at a fraction of the miss rate; the same input token therefore costs one thing the first time and another thing every time after, depending on whether the prefix hits. The printed rate is usually the miss rate, and the miss rate is what a workload pays if it never repeats a prefix.

Who controls the hit is the part a comparison tends to miss, and two design families show the difference. In one, prefix caching is automatic and the discount applies whenever the provider decides a prefix is reusable (OpenAI's automatic prefix caching); in the other, caching is an explicit, priced write with a lifetime, so a workload pays to create the cache and the discount only lands if the prefix is re-read inside the stated interval (Claude's explicit cache writes). The two cards can print a similar cached rate and behave differently under the same traffic: the automatic one rewards stability for free, the explicit one rewards stability only when the reuse is dense enough to amortise the write.

The consequence for a cheapest ai api question is direct. Until the mechanisms match, a ranking is meaningless: two vendors with the same miss rate are not the same vendor once cache reuse is in play, and a workload with no repeated prefix pays the worst rate on either card. The practical first measurement is therefore not a price at all — it is how much of your own context repeats.

cheapest llm: the off-peak window and the family tiers

The cheapest llm for a named task is not the smallest number on a card; on a low tier it is the member of the family that fits the task, priced under the timing rule the task can accept.

The timing rule is the off-peak window. Several providers publish a lower rate for work submitted outside peak hours, and the mechanism is the same everywhere: the provider trades a scheduling constraint for a discount, running the job when its capacity would otherwise idle. It is a real saving for batch scoring, backfills, nightly indexing and evaluation runs, and it is worth nothing to a request that has to answer a user now. A card with an off-peak rate is therefore two products — an interactive one at the peak rate and a deferred one at the window rate — and a comparison that mixes them is comparing a workflow to a workload.

The family rule is the ladder. What a vendor sells as "the model" is usually a set of models at different points: a small, cheap member tuned for short outputs and a larger one that is slower and dearer but does more per token. The cheap member is not a degraded copy of the expensive one for every task; for extraction, routing, classification and short answers the small rung is not merely cheaper but sufficient, and choosing it is the single largest cost decision available on a low tier. The larger rung earns its rate only when the task's quality bar genuinely needs it.

Two workload shapes sit outside all of this, and they are where the low card stops being cheap. The first is long output: output is the dear direction on every tier, and a low input rate does nothing for a job that writes thousands of tokens. The second is high concurrency: a low price and a generous-looking quota still meet a per-key rate limit, and once a workload is throttled the cost reappears as retries, queued latency and a larger context re-sent after each stall. A provider running on specialised inference hardware can push the per-token rate down on the models it serves and change the throughput ceiling at the same time (a provider running on specialised inference hardware), which is the reminder that price per token and price per completed task are not the same number.

free llm api: a zero rate with a ceiling

A free llm api is not a price of zero; it is a price of zero plus a quota, and the quota is where the cost actually lives. The rate is zero, but the shape is the same four dials as a paid tier, with the walls set hard: a monthly token allowance, a requests-per-minute limit, a cap on concurrent calls, sometimes a smaller model or a shorter context than the paid service. A workload that fits inside those walls pays nothing and is genuinely cheap; a workload that does not fit pays in the ways a bill never shows — a run that exhausts its allowance mid-batch, a retry loop that burns the handful of requests a minute allows, a queue that has to wait for the next window.

That is the honest way to read the free tier: compare it to the ceiling, not to the rate. A free tier with a token allowance your nightly job clears every night is not a free option; it is an option with a hidden monthly cost measured in failures. A free tier whose limits your workload never approaches is the cheapest option available, and no paid tier can beat a zero rate on the work that fits. The provider landscape itself — which services offer a free allowance today, and what each one caps — belongs to the sibling page on the free-tier LLM APIs themselves; the shape point here is narrower and does not change with the provider: a zero rate is cheap only up to the ceiling, and the ceiling is the real product.

One consequence follows for the low-price tier as a whole: a free allowance and a paid low tier are often the same model at the same underlying rate with a different wall around it, so the two answer the same question from opposite ends — one asks how much will fit under zero, the other how little the overflow costs.

llm token pricing: the unit decides what is metered

Llm token pricing is a convention as much as a figure, and the unit is the tell. A rate quoted per million tokens says the vendor meters usage and expects the number to be large; the same rate quoted per thousand says the same thing in a different unit; a price quoted per request or per seat says the vendor has chosen not to meter tokens at all. On a low tier the unit matters more than usual, because the lowness is concentrated in one class of token — the input side — and only a unit that separates the directions lets a reader see it. A single blended figure per million tokens has someone else's input-and-output mix baked into it, and it hides exactly the asymmetry a low tier is built on.

There is a second reason to read the unit carefully on a cheap card: the meter counts tokens, not characters and not words, and tokenisers disagree. The same paragraph can be a different number of tokens under two providers' encodings, and the disagreement is worst on precisely the text a cheap tier is often asked to handle — code, identifiers, non-Latin scripts. A rate comparison that assumes an identical token count across providers is therefore only approximately right, and the approximation is the difference between a low rate and a low rate on a different tokenisation.

cost per token: how our own meter records a low-rate model

A cost per token is only as trustworthy as the table behind it, so it is worth showing one in full — our own. The table is a local JSON file under the budget-guard module of the product source, read read-only from its origin branch on 2026-10-08. It holds four values per model: input cost per token, output cost per token, maximum input tokens and maximum output tokens. The unit is a single token, never "per 1K", so no implicit division sits between the stored rate and the arithmetic that consumes it. At that revision the row for deepseek-chat read:

Field Value
input cost per token 1.4e-7 USD
output cost per token 2.8e-7 USD
max input tokens 65,536
max output tokens 8,192

Three properties of that row are the ones a reader should carry, whatever the numbers are on the day they read this. First, the direction asymmetry is preserved in the table rather than flattened: the output rate is twice the input rate, so a per-million blended figure would hide the one property the shape of the tier turns on. Second, the context ceiling is a column, not a footnote — 65,536 input tokens and 8,192 output tokens are two different ceilings, and the smaller one is the harder limit for a generation-heavy job. Third, the table is a snapshot, not a quote: it is a diffable file that changes by commit, so it is only as current as its last edit, and any budget built on it should re-read the file rather than trust this page.

How usage becomes cost is the second half, and it is a two-part path. The central usage recorder keeps one monthly counter per team and increments it from the tools it treats as billable — compression, fetch, search and de-duplication — while explicitly skipping the budget check, the budget count and the rate-limit tools, which are metering calls rather than metered work. When a module reports a token count, that number is what the counter records. When it does not, the recorder falls back to an estimate: the character length of the returned content divided by four, with a floor of one, so a fetcher or a searcher that cannot report tokens still moves the counter instead of booking zero. The cost of a call is then that recorded count against the per-token values in the table above.

Two honest notes follow from reading it that way. The character-based fallback is an approximation of a token count by a fixed ratio, which is fine for budgeting and wrong for billing — it drifts on text that tokenises unevenly. And because the counter records tokens per tool rather than per model, the table's per-model rates apply where a model is named, and the shape of the tier is only visible when one is. The reading here is the shape of our own meter, not a bill: no arithmetic is reproduced, and the values above carry the date they were read.

Where a platform sits next to a low-price tier

A per-token low tier and a platform are different layers, and the distinction matters precisely for the workload this page has been describing. A cheap per-token tier prices the model call; it says nothing about the tool traffic that fills the prompt before the call, which on an agent loop is often the larger and the more variable part of the spend. A platform can price that layer directly.

That is where our own product sits. It does not sell tokens and does not fit the per-token row at all; it prices a platform over the tool traffic that shapes an agent's context, and it publishes operational limits rather than rates. The four tiers are defined by ceilings rather than by a rate: monthly token caps of 2M, 20M, 100M and 200M-plus; MCP requests per minute per key of 120, 300, 600 and 1200; audit-log retention of 7, 30, 90 and 180 days; and up to 2, 10, 30 and unlimited keys per team. Read against the four dials above, those columns are not prices — they are the ceilings and retention windows that decide what a workload is allowed to do, in the same way a free tier's quota is.

The billing model is the part worth stating plainly, because it is the difference between a platform and a markup: a platform fee, with a share of the saving only once the platform has measurably saved something, and the total capped at the top of a band. That structure does not grow with every call the way a per-token markup does, which is the mechanism a low-price tier and a platform answer differently. The caps, per-key rates, retention windows and key counts change by tier, so the pricing page is the authoritative table and this one is not.

How to get started

The first move costs nothing and is the one that prevents the rest: measure your workload against the four dials before you read a single rate.

  1. Measure cache reuse. Count how much of your context repeats between calls. On a low tier this single number decides whether you pay the hit rate or the miss rate, and it is a property of your workload, not of any card.
  2. Separate the directions. Split your expected usage into input and output tokens. A low input rate and a long-output job is the classic mismatch on a cheap tier.
  3. Decide whether timing is flexible. If the work can wait, an off-peak or batch window is a real discount; if it cannot, read the peak rate and ignore the window.
  4. Pick the family rung deliberately. Start with the smallest member of the family that clears the task's quality bar, and move up only when a measured result says to.
  5. Check the ceiling before the rate. For a free or low tier the allowance and the per-key limit decide whether the price is reachable at all.
  6. If your context is built from tool traffic, start free and watch one loop end to end, then confirm which tier your real volume needs.

Frequently Asked Questions

Is DeepSeek pricing actually the cheapest option?

It depends on the workload shape, not on the headline rate. A low input rate wins for input-heavy, cache-reusing, short-output work, and loses for long-output or high-concurrency work where the dear direction or the rate limit dominates. The card is a contract for a particular shape, so the real question is which shape your workload has.

Why is output priced above input on a low tier?

Because the two directions cost the provider differently. Reading a prompt is a parallel operation and generating is sequential, one token appended at a time, so a low card can hold the input rate down while leaving the output rate relatively high. That asymmetry is why a low input rate and a long-output job are a mismatch.

What does an off-peak rate actually trade?

Latency for money. The provider accepts your work into a window it can schedule on otherwise idle capacity and passes part of that saving back as a lower rate, so the discount is real only for jobs that can wait. A request that must answer immediately never sees it.

Does a cache hit make input free?

No, it makes it cheaper, and only on the tokens that hit. A kept prefix is re-read at a fraction of the miss rate, so the effective input price depends on how much of your context repeats and on whether the provider caches automatically or requires an explicit, priced write. A prefix that never repeats pays the miss rate every time.

When is a free tier not actually free?

When the workload runs into its ceiling. A zero rate with a token allowance, a per-minute request limit and a concurrency cap is cheap only while the work fits underneath them; past that, the cost reappears as failed runs, retries and queued latency rather than as a line on a bill.

Limitations

This page is a shape, not a price list, and it deliberately reproduces no provider rate it did not re-check on the day it was read; a card always wins over any summary, including this one. The four dials are described generically, and real vendors combine them in ways a four-row description only approximates, so the ranking of any two providers depends on your own usage mix rather than on anything written here. The page also stops short of the arithmetic: it explains what turns a low rate into a low cost and where that stops being true, but it does not compute a bill, does not model one specific request, and does not price self-hosted hardware. Where a number would have moved, the page names the mechanism instead and points at the provider's own table.

Sources

Method note

This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 1 of 7 sections for this page (0 abstentions, 6 no-slice verdicts): the single pin is a general-purpose costing helper extracted from a third-party library, and its inline documentation is not written in English, so quoting it verbatim would place non-English text on an English page for no evidential gain. The section above that carries our own reading instead reports the price table that helper consumes, read read-only from the product source, which is the same first-hand evidence in a form this page can use. Every other section is written from the providers' own public documentation, cited with a checked date, because the house rule for a section without a usable pin is sourced, never invented. Batch fingerprints live in this project's pipeline_results.json and the spec docstring, never on this page.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)