SmartGate

Context Compression: Four Methods and the Fidelity Each Costs

Context compression is not one technique but four different operations: summarization rewrites a passage in fewer tokens, extraction keeps the highest-scoring spans verbatim, key-value structuring moves text into typed fields, and de-duplication collapses repeated passages into one. Each discards a different kind of information, so a compression ratio on its own proves nothing.

Short answer: Context compression is not one technique but four different operations: summarization rewrites a passage in fewer tokens, extraction keeps the highest-scoring spans verbatim, key-value structuring moves text into typed fields, and de-duplication collapses repeated passages into one. Each discards a different kind of information, so a compression ratio on its own proves nothing. The measurement that does is the ratio reported next to the task success rate on a fixed set of cases.

Key takeaways

  • A ratio is a cost number, not a quality number. A five-fold reduction means nothing until you know which facts left with it.
  • Recoverability separates the four methods. Extraction and de-duplication keep verbatim spans that a retained source can restore; summarization and key-value structuring produce a new artefact that cannot be inverted.
  • Compression earns its cost only when the budget is actually exceeded. Shrinking a call that already fits spends latency and buys nothing.
  • Where compression runs — before the call, between steps, or before a write — decides which failures you can still detect afterwards.
  • Report three numbers together: tokens saved, task success against an uncompressed baseline, and the tokens the compression step itself spent.

context compression: four operations and the fidelity each one costs

Context compression is any step that reduces the tokens a model receives while keeping whatever the task needs in order to answer correctly. The definition is deliberately about the task rather than about size. A reduction that is ninety percent smaller and loses the one figure the question asked for is not compression, it is deletion with a better name. This page treats the four operations below as distinct because they fail differently, and a team that has only one of them in its code has only one failure mode to test for.

Summarization rewrites the material. A language model — sometimes the same one, sometimes a cheaper one — reads a passage and produces a shorter version in its own words. The fidelity loss is interpretive: values get rounded, a negation can vanish, an attribution ("per the 2023 policy") can be dropped or reattached to the wrong claim, and the ordering that carried a causal argument gets flattened into a list. Because the output is fluent, none of this is visible without a check, which makes summarization the method that most needs an evaluation set rather than a reading.

Extraction keeps selected spans. A scorer ranks sentences, chunks or tokens and the top share is passed through unchanged. LLMLingua is the token-level shape of this idea and Selective Context is the sentence-level one. The loss is coverage rather than wording: the model receives real sentences, so it cannot hallucinate inside them, but the sentence carrying an exception or a condition is exactly the kind a relevance scorer tends to discard, especially when the exception reads as a tangent.

Key-value structuring moves text into typed fields. Dates, owners, amounts, statuses and identifiers are lifted out of prose and into a schema. The loss is whatever the schema has no field for, plus any ambiguity the schema forces into a single slot. When the schema matches the source, precision is high and downstream code can validate it; when it does not, a value lands in the wrong field and the result is grammatically perfect and wrong. This is the method whose failures are easiest to test — a schema is a test — and easiest to miss, because nothing about the output looks broken.

De-duplication collapses near-identical passages into one representative. The loss is subtler than the other three: repetition itself carries signal, because a fact repeated across three documents is usually more load-bearing than one mentioned once, and near-duplicates are rarely identical. The second copy is often where someone added the caveat that the first one omitted. Keep one representative and you keep the topic; you do not keep the emphasis.

The four operations answer one question differently: can the original be re-derived? Extraction and de-duplication hand the model real spans and can therefore be undone as long as the source is still somewhere; summarization and key-value structuring hand it a derivative, and a derivative is a new artefact with its own errors. That distinction is what decides where a compressed copy is allowed to live. The mechanics here are one half of context engineering, which owns the wider question of what belongs in a model call at all. The other half, deciding what to fetch in the first place, is a retrieval problem.

Operation What it keeps What it loses Invertible from the compressed form alone
Summarization the gist, in the compressor's own words exact values, negation, attribution, order no
Extractive selection the spans the scorer ranked highest every unselected span, conditions included only from a retained source
Key-value structuring the fields the schema names everything outside the schema, ambiguity collapsed no
De-duplication one copy of a repeated passage that it repeated, and how the copies differed only from a retained source

llm memory: what a compression pass is allowed to drop

Memory and compression are usually discussed as one topic and they are not the same operation. Memory is a durable store with a retention rule and an owner; compression is a per-call reduction with a token budget. They meet at one decision: whether the compressed copy is allowed to become the only copy. The house rule is that it is not. A store keeps the source, and reduction happens on the read path, at the moment a context is assembled. Compress on the write path and your retention setting stops meaning anything, because the oldest detail is gone on day one rather than on day 180.

Once the source is safe, the question becomes what a reduction may strip. Four classes of content should never be dropped, whichever method runs:

  • Identifiers. Account numbers, commit references, file paths, ticket ids. A summary that replaces an identifier with "the related record" has destroyed the only thing that made the sentence verifiable.
  • Negation and modality. "Not yet approved" and "approved" differ by one token and collapse to the same summary of the subject matter. Modals matter for the same reason: "may write" and "must write" reduce to the same verb in a short paraphrase.
  • Units and precision. A figure without its unit is a different figure, and rounding is a decision that belongs to the reader, not to the compressor.
  • Provenance and recency. Which document said it, and when. A memory that keeps a fact but not its source cannot be re-checked, and a fact from last year reads identically to one from this morning.

A memory layer with those four rules written down can accept aggressive reduction on everything else, because the exceptions are the ones that break runs. What is left is ordinary store design: what a write is allowed to contain, who reviews it, and how long it survives. The durable half of this picture — what memory is, what may be written into it, and what is read back on the next turn — is the subject of llm memory; this page only claims the reduction side.

memory agent: three insertion points and their failure modes

Where a memory agent compresses decides what it can still detect. There are three insertion points, and each needs a different control, so choosing one is a design decision rather than a performance knob.

Before the call. The working context for one model call is assembled and then reduced. This is the cheapest place to compress and the hardest place to audit, because the drop is per-call and leaves nothing behind. The control is a baseline plus an evaluation set: run the same cases with compression off, and treat the difference as the cost of the reduction. Without that baseline, the failure is invisible until a user reports a wrong answer that the model could not have avoided.

Between steps. A loop has run twenty steps and the history no longer fits. Compressing the history mid-loop is necessary and dangerous for one reason: the summary becomes the agent's only record of what it already tried. An agent that loses its own transcript repeats a step it already completed, or worse, treats a failed attempt as unfinished work. The control is a verbatim step log that lives outside the summary — the reduction feeds the next prompt, and the raw steps stay in a table.

Before a write. The agent decides what deserves to enter long-term memory, and de-duplicates or derives it before the write. Here the risk is category confusion: an interpretation becomes a stored fact the moment it is written. The control is review on the write path, not on the read path, because a bad read can be retried and a bad write is quoted back with full confidence on every later run.

The three points are not alternatives; a serious agent does all three and keeps three different records. The agent that reads and writes the store, and the loop those writes happen inside, are the subject of memory agent. This page's contribution is narrower: name the insertion point, decide the control, and log the compression event so the reduction is attributable after the fact rather than assumed.

agent memory systems: keep two artefacts and price the reduction

Read as a system, a memory layer with compression in it holds two artefacts, not one. The first is the source of truth — the transcript, the document, the tool output as it arrived. The second is the working copy — whatever reduced form the next call actually receives. The working copy may be rebuilt at any time from the source, and it must never overwrite it. Teams that skip that rule end up with a store whose oldest entries are one-sentence summaries of documents nobody kept, and no way to recover the values those sentences left out.

Two operational questions follow, and they are boring and decisive. First, does each working copy have a lifetime, an owner and a rule for what happens when it expires? A derived artefact with no expiry becomes a second source of truth by accident. Second, does the compression step pay for itself? A reduction that saves tokens on the model call but spends tokens on a compressor is only net positive past a break-even point, and the break-even depends on how often the compressed copy is reused. Compressing once and reusing the result across many calls is a different economics from compressing on every call.

The metric that survives budget review is not tokens saved but cost per successful task: total tokens and latency, compressor included, divided by the number of tasks that produced a correct result. A compression layer that halves the prompt and drops the success rate by a third has made things worse, and the tokens-saved figure will never show it. That framing — total cost against verified outcome — is the same one a quota or budget control needs, which is why the two are usually built together.

url to markdown: the first reduction is a format change

Before any of the four operations above can run, a fetched page has to become text. Converting an HTML document to markdown is compression in the broad sense and the cheapest kind of it: scripts, styles, navigation, cookie banners and tracking markup are removed and the visible text is kept. Unlike summarization it is deterministic and, in structure, largely invertible — the headings and table rows are still there in order. Its fidelity loss is precise and easy to under-rate: exactly the non-text information goes. Tables survive as pipe-separated rows, images reduce to whatever alt text someone wrote, and meaning that lived in layout rather than in a sentence is simply absent from the result. A conversion step that drops a pricing table's column header has compressed nothing and corrupted everything downstream of it.

The tooling around that step splits into two families: open converters that take a URL and return markdown, and managed fetch-and-convert services that add rendering, retries and extraction policy. The conversion itself, and the differences between the free converters, is the subject of url to markdown; the managed end of the market, where a service decides what to keep before you ever see the text, is compared on Firecrawl alternative.

This page carries one code excerpt, and it is worth being exact about what it is. The matcher that pins excerpts to section keywords found a single unique symbol for the phrase this section is built around, and that symbol is a keyboard-shortcut handler in the product's dashboard command palette rather than a converter. It is kept because the house rule is that excerpts are cut from the repository and asserted verbatim, never typed by hand — and because it happens to mark the boundary this section is about: the handler below opens the client surface where a URL is entered, and everything after that line, the fetch, the format change and every reduction that follows, happens on the server.

# components/dashboard/search-command.tsx — source lines 24–29 (the symbol name)
const down = (e: KeyboardEvent) => {
      if (e.key === "k" && (e.metaKey || e.ctrlKey)) {
        e.preventDefault();
        setOpen((open) => !open);
      }
    };

Read that for what it is: six lines of client code with one job, opening a dialog on a keyboard shortcut. It is the smallest possible example of a reduction that is not really a reduction at all — a UI affordance rather than a transformation of content — and that is precisely why it belongs here. The moment a URL leaves that dialog, the page is fetched and converted by code this excerpt does not show, and the fidelity question from earlier in this page is already live. Nothing in the excerpt compresses anything, and the honest reading is that the pin is a weak one; that finding is recorded in the Method note rather than dressed up.

agentic retrieval: retrieve less instead of compressing more

Compression and retrieval both reduce what a model reads, and they are substitutes more often than a team admits. Retrieval reduces the token count by fetching fewer, better passages; compression reduces it by shrinking what was already fetched. Reranking a candidate set is extraction at the document level — the same coverage-versus-wording trade, one layer up.

That framing gives a clean rule. Retrieval wins when the corpus is large and the query is narrow: five relevant passages out of two hundred index entries should be found, not compressed from two hundred. Compression wins when the material has already been selected and all of it is needed, but it is longer than the budget: a long conversation, a large tool output, a document the question is specifically about. The two are not in tension in that case, because there is nothing left to filter.

The failure the distinction catches is a team compressing an over-large context instead of fixing an over-broad query. That pays twice — a loose retrieval pass loads the window, and a compressor then trims what should never have been loaded — and it hides the real defect, which is that the query asked for too much. Before adding a compressor, the cheaper experiment is to halve the number of results and rerun the case set. If success holds, the problem was retrieval and the compressor was about to mask it. The loop that fetches, ranks and decides when to stop is the subject of agentic retrieval; this page only claims the decision of whether a compressor belongs in front of it.

prompt compression: report the ratio with the success rate

Prompt compression is the token-level branch of the same family, and it is where the measurement discipline has to be stated most carefully, because the headline number is so easy to quote. The compression ratio is compressed tokens divided by original tokens — always state the direction, since "a ratio of 5" describes two different results depending on which way it is written. Tokens saved is the subtraction. Published work in this branch reaches large reductions: LLMLingua reports up to a twenty-fold compression with a small quality drop on its benchmarks, and gist-token methods report similar orders of magnitude. Those results are real and they are also benchmark-specific, which is why the number cannot be carried across to your own traffic without a check.

The check is a paired evaluation, and it is short:

  1. Freeze a case set where each case has something checkable in the output — a required value, a schema, a refusal, or a tool call that must not happen.
  2. Run the whole pipeline on that set with compression off. Record the success rate, the tokens and the latency.
  3. Run the same set with compression on and the same revision otherwise. Record the same three numbers.
  4. Compare. If success holds within the tolerance you set, the reduction is free; if it does not, the next step is to localise the loss to one of the four operations rather than to "compression".

Localisation is the part teams skip, and it is the part that makes the result usable. Plant a fact in each case that the answer must use — a figure, a negation, an identifier — and measure how often it survives. Summarization leaks values; extraction leaks coverage; key-value structuring leaks anything the schema does not name; de-duplication leaks the second copy's caveat. Testing for the loss that matches your method turns a single pass-fail number into a fix.

Metric Definition What it tells you
Compression ratio compressed tokens divided by original tokens the per-call cost, and nothing about quality
Tokens per successful task all tokens, compressor included, divided by successes whether the saving is net
Task success rate share of cases whose checkable property holds whether the output is still right
Lost-fact recall share of the planted facts present in the answer which operation lost what
Added latency compressor time per call the cost the ratio hides

A ratio reported without the other four is a sales figure. The four together are a decision: they say whether to keep the compressor, which operation to keep, and where it is allowed to run.

How SmartGate fits

SmartGate is a gateway in front of model calls, which is the natural place to apply a reduction once rather than in every application that talks to a model. The plan table sets the budgets those reductions compete against: monthly token caps are 2M, 20M, 100M and 200M+, requests per minute per key are 120, 300, 600 and 1200, audit-log retention is 7, 30, 90 or 180 days, and a team can hold 2, 10, 30 or 9999 keys. Read against this page, one row matters more than the others. A compressed artefact inherits the retention of the record it came from, so if a summary is the only thing left after the source's retention window closes, the audit trail cannot tell you which call dropped the value. The audit log is what makes a compression event attributable after the fact, and its length is a plan decision rather than a technical detail. The pricing page is the authoritative table for those figures, and it is the one to read before committing a retention requirement to a date.

Frequently Asked Questions

Limitations

This page describes four operations and a way to measure them; it does not rank compression libraries, and it quotes no benchmark of its own. The published figures it refers to come from the papers listed below and belong to their benchmarks, not to any particular production traffic — the only number that should drive a decision is the one measured on your own case set.

It also does not claim that every context should be compressed. Three of the sections argue the opposite at least as often: a call that fits should be left alone, a narrow query should be retrieved better rather than compressed, and a compressed artefact should never become the only copy of its source. The recoverability split is a design rule, not a guarantee, and it depends on the source actually being retained under a retention policy someone owns.

Finally, this page carries one code excerpt, and the Method note below records why there is only one and why it is weak. Readers should not read a single pinned symbol as evidence about how any converter, gateway or memory store is implemented; it is a verbatim fragment of client code and nothing more.

Sources

  • LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models — arxiv.org/abs/2310.05736, for token-level prompt compression, the budget controller idea and the reported twenty-fold compression with a small quality drop.
  • Lost in the Middle: How Language Models Use Long Contexts — arxiv.org/abs/2307.03172, for the position effect that makes what a reduction keeps and where it puts it a quality question rather than a pure cost question.
  • Learning to Compress Prompts with Gist Tokens — arxiv.org/abs/2304.08467, for the activation-level branch of prompt compression and its reported compression rates.
  • Demand figures in this page are our own measurements: DataForSEO Google Ads, United States, 12-month window, measured 2026-10-01, recorded in this project's search_volume.json and research_brief.md.
  • Product behaviour and the plan table: read from the product source at the revision pinned in this project's pipeline_results.json, read-only, with the plan figures re-verified against the live pricing page on 2026-10-01.

Method note

The single fenced block on this page was cut out of the slice body returned by the API and re-asserted against it byte-for-byte, never typed by hand; the first line inside the fence records the file and the source lines it came from, and the table below gives the row for it.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)
1 url to markdown down components/dashboard/search-command.tsx 24–29 rule A L2 → slot-proof 42939cf66801

This page carries one code excerpt, and the honest reading of it is that the pin is weak. The slice matcher pinned 1 of 7 sections for this page (0 abstention(s), 6 no-slice verdict(s)): only the section keyed to "url to markdown" resolved to a unique local match, at rule A level 2, and the symbol it found is a keyboard-shortcut handler in the product's dashboard command palette rather than anything that converts or compresses content. The other six section keywords returned no-slice, because this neighbourhood's vocabulary collides with generic helper and type names across a codebase.

The excerpt is still cut from the repository and asserted verbatim rather than typed by hand, which is the rule that keeps a weak pin honest: the page shows what was found and says plainly what it is. Product claims were read from the product source at the revision the slice run recorded in this project's pipeline_results.json, read-only, and the plan figures were re-verified against the live pricing page on 2026-10-01. The section keyword quoted above each heading comes from this project's own paid measurement run, not from a third-party tool. No code, batch fingerprints, auction data or internal hosts are transcribed, so nothing here has to be asserted verbatim beyond the single block above.