Free LLM API: How to Read a Free Tier Before You Commit
A free LLM API is not a price of zero; it is a set of constraints you agree to operate inside. Four of them decide almost everything — a request or token rate limit, a quota or trial credit with an expiry, a model subset that leaves out the largest and newest models, and data-use and retention terms. Judge a free tier by how cheaply you can leave it, not by how many model names it lists.
Short answer: A free LLM API is not a price of zero; it is a set of constraints you agree to operate inside. Four of them decide almost everything — a request or token rate limit, a quota or trial credit with an expiry, a model subset that leaves out the largest and newest models, and data-use and retention terms. Judge a free tier by how cheaply you can leave it, not by how many model names it lists.
Key takeaways
- "Free" describes a boundary, not a price. The number worth comparing is what happens when you cross each ceiling.
- The first limit you hit defines the tier. An interactive app meets requests per minute first; a batch job meets the quota, or its expiry, first.
- A free tier is a trial of the interface, not only of the models. The auth shape, the SDK and the request format are what you carry away.
- Switching cost is a real price. A tier on an OpenAI-compatible protocol is cheap to leave; one with a bespoke SDK, a scoped key and server-side state is not.
- Data terms are limit clauses too. Whether prompts train a model, how long they are kept and who may read them belong beside the rate limit.
- Do this first: score the candidate on the checklist at the end of this page — rate limit, quota, model subset, terms, exit cost — before writing integration code.
Most teams meet their first LLM API through a free tier, and many later discover that "free" described a boundary rather than a price. The discovery is unglamorous: a demo that answered instantly in testing starts returning rate-limit errors in production, a trial credit that felt unlimited runs out three weeks in, or a model that worked in the prototype is quietly not in the free subset. None of those is a defect. Each is a clause of a contract the word "free" never mentioned.
This page treats a free tier as what it is — a small structure of limits and terms — and gives a way to read it before committing. It is deliberately not a ranking of free providers or a price list: the two providers whose rate cards already have their own comparison pages own the "cheapest" question. The narrower question here is: given a free tier, what is it promising, what will it do when you grow past it, and how much will it cost to stop using it. The centre that maps how providers charge at all is how providers charge.
free ai api: the four things the word free hides
A free ai api is described by four axes, and a page that shows only one of them — usually the model names — has told you almost nothing. The four are the rate limit, the quota, the model subset and the terms. They fail in different ways, so they belong in separate columns of any comparison, even one you make only for yourself.
The rate limit is how fast you may call. It is usually written as requests per minute, sometimes as tokens per minute, and often per key rather than per account, which matters because a single key spreads its allowance across every process that shares it. A rate limit is the first ceiling an interactive app meets, and the one that most often turns a working demo into a flaky one: nothing changed except the concurrency the demo never had.
The quota is how much you may call in total. It takes three common forms — a monthly allowance that resets, a one-time trial credit that does not, and a hard request or token count that simply ends. The form matters more than the number, because a resettable allowance and an expiring trial fail identically on the day they run out and differently in every budget model that tries to predict them.
The model subset is which models the free rate applies to. It rarely includes the largest or newest model, and it frequently excludes the long-context variants, so a workload that depended on a very long context window may not have one in the free tier at all; a provider that publishes an explicit long-context tier, such as Claude's cache and tier pricing, is at least stating that boundary. The subset is also where a free tier most easily disguises a paid upgrade as a bug fix: the model that "stopped working" was never in the free set.
The terms decide what happens to your data — whether prompts and completions are used to train models, how long they are retained, whether humans can review them, and where they are stored. These are limit clauses exactly as a rate limit is, and they are the ones a team is most likely to have accepted without reading, because they sit below the fold of a page whose headline said "free".
A fifth axis — how you authenticate and shape a request — is not a limit, but it decides the cost of leaving, so it gets its own section below. For now, the rule is that a free ai api is a structure of speed, volume, model access and terms: two of the four you meet in the first hour of load testing, and two in procurement, where they are harder to fix.
cheapest llm api: cheapness lives inside the boundary
A cheapest llm api search is usually a search for a rate card, and a free tier is the one rate card whose headline number is zero. The trouble is that zero is the least informative price there is, because the cost of a free tier is not in the requests it serves but in the requests it refuses. A free tier's effective price curve has three segments, and which one your workload lives in decides the only comparison that is real.
Below the limit, the marginal price of a request truly is zero, and no paid tier beats zero; for a prototype, an internal tool or a low-traffic feature, that segment is the whole story. At the limit, the marginal price of the next request is not money but latency and failure — the call is refused or queued, and your application has to handle it. Above the limit, the ways forward are to acquire a second key if the terms allow it (they often do not), to move up to a paid tier of the same provider, or to migrate elsewhere. The first is a terms question, the second a price question and the third a migration question, and a list of free endpoints answers none of them.
Ranking free tiers by "cheapest" therefore collapses into a question about your own load shape. A workload that peaks briefly is cheapest on whichever tier has the highest burst rate; one that runs steadily is cheapest on whichever has the largest total allowance; one that must never fail is not served by a free tier at all, and its cheapest option is the smallest paid tier that removes the refusal. That is why this page prints no leaderboard: the honest output of a comparison is a sentence about a named workload, and a single winner would hide the workload it was computed for. The two providers whose rate cards have their own comparison pages are exactly the cases where a list does make sense — DeepSeek's per-token price page and a marketplace's pass-through pricing — and neither is a free-tier question.
There is a further trap in the zero. A tier generous enough to run a prototype is often generous because the prototype is the funnel, and a funnel is designed to end. That is not dishonest, but it means the boundary is a product decision rather than an accident: the tier is shaped so that the moment your usage becomes valuable, it also becomes billable.
cheapest ai api: the costs a price list never shows
The cheapest ai api is not the one with the lowest number on it, because the costs that matter most never appear on a price list. For a free tier those are the whole question, since the money it saves is real only until the first of them arrives. Three are worth naming before any provider is compared.
The first is the integration cost of leaving. Every hour spent wiring a bespoke SDK, handling a bespoke error model and normalising a bespoke response shape is an hour that must be spent again elsewhere, or spent on an adapter that keeps two shapes alive at once. An interface that matches a widely supported request format is cheap to leave because the code does not have to change; one that is unique is expensive to leave for the same reason it was pleasant to adopt.
The second is the cost of the boundary itself. When a limit is hit the request does not simply stop; the system around it must decide whether to retry, downgrade, queue or fail. Engineering that behaviour is part of the price of a free tier, and it is paid in code rather than in dollars. A tier with a small, well-understood limit is often cheaper overall than a larger tier whose limit arrives rarely and therefore without a plan.
The third is the cost of being migrated for you. A free tier that ends, changes its model subset or alters its terms does not wait for your roadmap. Each of those is a small migration you did not schedule, and the more your product depends on the exact free behaviour — the exact context length, the exact model, the exact tolerance for latency — the larger that migration becomes.
Together, those three say something a price list cannot: the cheapest ai api for a given workload is the one whose boundary you can predict, whose exit you can afford, and whose behaviour does not change under you.
cheapest llm: four ways a free tier holds you
"Cheapest llm" and "easiest to leave" are not the same claim, and the gap between them is where lock-in lives. A free tier holds a team in one of four ways, and it is worth checking which apply before the free rate has quietly become a dependency.
The first is protocol shape. If the tier accepts a request format other providers also accept — the shape OpenAI's own rate card popularised — migration is mostly a key swap and a base-URL change; if it requires a bespoke format, migration means rewriting the call site. The protocol, not the price, is the single largest determinant of exit cost.
The second is authentication and key shape. A single bearer key that drops into an existing client is cheap to move. A scheme that binds a key to an account or a project, rotates it, or scopes it to particular models adds work to every move. The question is not "is the key secret" but "what else does the key carry", because whatever it carries is something a migration must reconstruct.
The third is data and state. If prompts, completions, embeddings or conversation history live inside the provider, they are asset gravity: leaving means exporting, re-indexing or reconstructing them. A stateless tier is cheaper to leave than one that has become a store, even when the store is convenient.
The fourth is the cliff. A limit generous enough to disappear from daily awareness arrives as an outage rather than a signal. A tier whose limits are visible early trains the team to see the edge; one whose limits are hidden until they are crossed converts a budget question into an incident.
None of these is a reason to avoid a free tier. They are the reasons to choose one knowingly. Before adoption, write down what a migration out would cost in each of the four — protocol, key, state, cliff — and treat a tier that scores cheaply on all four as free in the way that matters, and a tier that scores expensively as a rental whose first month happened to cost nothing.
llm api pricing: how to read a free tier's terms
Reading llm api pricing for a free tier differs from reading it for a paid one, because the published numbers are limits rather than rates, and the clauses beside them carry as much weight as the figures. A short probe sequence turns a marketing page into a usable specification.
Start with what is limited and how. Find the rate limit (requests or tokens per minute, per key or per account), then the total allowance (monthly or one-time), then how a breach is signalled — an HTTP status, a header, or silence followed by failure. A limit you can observe in a header is one you can schedule around; a limit you learn about only from errors is not.
Then find which models and context sizes the free rate covers, and whether the subset is stable or expected to change. A free rate that applies only to a model the provider has already announced it will retire is a temporary tier, whatever the page implies.
Then read the data-use and retention clauses as limits in their own right. Four questions are usually answerable: are inputs or outputs used to improve or train models, how long are they retained, can anyone review them, and where are they stored or transferred. If a page answers none of the four, the honest reading is that the answer is not published — not that the answer is favourable.
Then find the overage behaviour: at the boundary, does the tier stop, throttle, queue, or charge. A free tier that stops is predictable; one that silently converts to metered billing is a free tier with a paid tier behind a door you did not open.
Finally, confirm the exit facts — the protocol shape, the auth shape and whether any state lives server-side — which the previous section flagged as the lock-in checks. A page that answers all of those is a tier you can adopt deliberately; a page that answers none is one whose limits will introduce themselves later.
ai api pricing: the self-check, and the limits our own gateway enforces
The checklist below is the one this page recommends, and the honest way to introduce it is to apply it to our own product first. SmartGate is not a model provider and does not resell tokens; it is a metering and control layer over the tool traffic an agent produces before a model is asked to complete anything. What it offers as a free tier is therefore a rate-limit structure rather than a model catalogue, and that structure is readable in the source our own limiter uses.
Read on 2026-10-08 from the plan entitlements module, the catalogue defaults are:
| Plan (catalogue default) | REST write requests/min | MCP requests/min per key | MCP requests/min per team |
|---|---|---|---|
| Free | 20 | 30 | 30 |
| Pro | 120 | 300 | 600 |
| Teams | 300 | 600 | 3000 |
| Enterprise | 600 | 1200 | 9999 |
Three properties of that table say more than the numbers. First, it is a catalogue default: the merge step lets a team's stored feature overrides replace any row, so the limit a specific key actually sees can be higher. Second, the resolver fails safe: when the lookup for a team fails, or when no team id is present, it returns the Free set — the entry tier is literally the fallback path, not a special case bolted on beside it. Third, the result is cached in the process for about a minute, so a plan change is visible within a short window rather than instantly. The monthly token ceilings that pair with these rates are the published plan caps (two million, twenty million, one hundred million, and two hundred million-plus tokens per month) and they move with plan changes, so our pricing page is the authoritative table and this page is not.
None of that is presented as anybody else's free tier. It is our own limit structure, quoted so the checklist has a worked example: a free tier is a rate limit, a ceiling that can be overridden per account, a fallback, and a cache interval. Whether the same is true of the tier in front of you is the point of the self-check:
- What is the first limit you will hit? Name it — a rate, a volume, or an expiry — and say which workload meets it first.
- Is the limit per key or per account? A per-key limit shared across processes behaves like a much smaller limit than the printed number.
- What is the quota, and when does it end? Distinguish a resettable allowance from a trial credit; the second is a countdown.
- Which models and context sizes are in the subset? Confirm the exact model and window your workload needs are inside it, not merely nearby.
- What do the data-use and retention terms say? Get answers to training, retention, review and storage, or record that the tier does not publish them.
- What does leaving cost? Write down the protocol shape, the auth shape and any server-side state, because those are the exit price.
Score the candidate against all six before writing integration code, and again after the first real load. A tier that answers all six is one you can adopt on purpose; one that cannot answer two is one whose limits you will meet the hard way.
How a metering layer's free plan differs from a provider's
The free tiers you will compare do not all cap the same thing, and a short comparison makes the difference concrete. The rows are the four axes from the first section; the columns are the three shapes a free tier usually takes.
| Axis | A provider's own free tier | A marketplace's free credits | A metering layer's free plan |
|---|---|---|---|
| Rate limit | requests or tokens per minute at the provider | per key, shared across the pooled models | requests per minute over the layer's tools |
| Quota | a monthly allowance or a trial credit | credits that expire, often quickly | a monthly platform allowance |
| Model subset | the cheaper and older models | whatever the pool routes you to | not a model question at all |
| Terms | the provider's training and retention clauses | the marketplace's, plus each provider's | the layer's own data handling |
| Where it hurts | the model you needed is excluded | the credits run out mid-build | the tool traffic, not the completion |
Read that way, a free tier is a shape, not a bargain. A provider's free tier is cheap on the models it wants to promote, and one running its own inference hardware can price that promotion lower (Groq's inference-hardware rates); a marketplace's credits are cheap until they expire; a metering layer's free plan does not touch the model at all, which is why a team can often run a provider's paid model behind a layer's free plan and be metered only on the traffic the agent generates. The comparison is not "which is cheapest" but "which ceiling lands on the thing I depend on", and a layer that meters tool traffic rather than completions caps a different resource by design.
How to get started
The first move costs nothing and prevents the mistakes the rest of this page describes: read the tier before you trust it.
- Pick one candidate and answer the six self-check questions on paper, using its own live documentation rather than a roundup post.
- Load-test the rate limit with the concurrency your real feature will have, not the concurrency a demo has.
- Write the exit note — protocol, auth and state — while the tier is still new and cheap to leave.
- Decide the boundary behaviour in advance: retry, queue, downgrade or fail, chosen per request path rather than discovered in production.
- If a layer will sit between the key and the model, start free so the tool traffic is metered from the first call, then confirm which tier your real volume needs on the pricing page.
Frequently Asked Questions
Is a free LLM API really free?
It is free in the sense that no invoice arrives for the requests the tier accepts, which is a real saving at low volume. It is not free of a price: it is bounded by a rate limit, a quota, a model subset and data terms, and the cost appears when you cross one of those or when you have to leave. Treat the boundary, not the zero, as the thing you agreed to.
What is the difference between a rate limit and a quota?
A rate limit caps how fast you may call, usually requests or tokens per minute, and it decides whether your working application stays working under load. A quota caps how much you may call in total, either as a monthly allowance that resets or as a trial credit that expires. Your app meets the rate limit first; your budget meets the quota first.
Why does the model subset matter so much?
Because the free rate usually applies only to the cheaper or older models, and often excludes the long-context variants. A prototype that depended on a large context window or a newest model may find neither is in the free set, so the feature that seemed to work was running on something the tier does not actually offer at that price.
How do I judge whether a free tier will lock me in?
Check four things before adopting it: the protocol shape, the authentication and key shape, whether any data or state lives server-side, and how visible the limit is before you cross it. A tier using a widely supported request format, a single bearer key and no stored state is cheap to leave. One that adds a bespoke SDK, a scoped key and server-side state is not, whatever it costs.
Are data-use terms part of the price?
Yes. Whether inputs and outputs are used to train models, how long they are retained, whether humans can review them, and where they are stored are limit clauses in the same sense a rate limit is. A tier that answers all four is one you can assess; one that answers none is one whose terms you will meet later.
Limitations
This page is a method, not a survey, and it names no provider's free tier and quotes no price. That is deliberate: a free tier's limits and terms change without notice, so any list of current numbers would be stale by the time it was read, and the reliable part of the problem is the reading method rather than the figures. Where a specific provider's rate card is the subject, that is a sibling page's job, linked above.
The self-check is a reading aid, not a guarantee: it cannot tell you whether a tier will change its terms next quarter, and it does not substitute for the provider's own documentation, which always wins over any summary. The worked example is our own product's limit structure and is presented as such, not as a claim about any other tier, and the plan figures it references move with plan changes, so the pricing page remains authoritative.
Finally, the four lock-in mechanisms are described generically. Real tiers combine them in different proportions, and a team's tolerance for each depends on workloads this page cannot see. The right use of this page is to run the six questions against a candidate you already have in mind and let the answers, not this text, decide.
Sources
- OpenAI API documentation and terms of use, for the request-format conventions that make a tier cheap or expensive to leave: https://platform.openai.com/docs/ and https://openai.com/policies/terms-of-use
- Anthropic usage policies and API documentation, for the data-retention and training clauses a free tier may carry: https://www.anthropic.com/legal and https://docs.anthropic.com/
- Google Gemini API pricing and terms, for a worked example of a free rate that covers only a subset of models: https://ai.google.dev/gemini-api/docs/pricing
- Product behaviour and the plan ceilings: read read-only from the product source at the revision recorded in this project's pipeline_results.json, with the figures re-checked against the live pricing page on 2026-10-08.
- The self-check's worked example is our own implementation, read on 2026-10-08 from backend/smartgate/core/plan_entitlements.py (origin/main): the catalogue defaults, the per-team override merge, the fail-safe to the Free set, and the 60-second cache.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of the 7 planned sections: rule A found no unique symbol in the scanned repository for any section keyword, and the remote candidate fallback returned only score-ranked near-misses, which a page must not dress up as a pin. Every section above is therefore written from public sources — the providers' own policies and pricing pages, cited with a checked date — because the house rule for an unpinned section is sourced, never invented. The one exception is "ai api pricing: the self-check, and the limits our own gateway enforces", written from our own implementation: the plan and rate-limit catalogue read read-only on 2026-10-08 from backend/smartgate/core/plan_entitlements.py, stating the catalogue defaults, the per-team override merge, the fail-safe and the cache interval, and labelled throughout as our product's structure rather than any third party's free tier. The section keyword quoted above each heading is this project's own measured pool phrase, not a code symbol, and every one of them carries a measured search volume above zero. No batch fingerprints, auction data or internal hosts appear in the text.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|