SmartGate

Context Engineering: Four Mechanisms That Fill a Window

Context engineering is the practice of deciding what a model sees on each call, and it replaced prompt engineering because the costly failures moved from wording to supply — which documents retrieval returns, what compression drops, what memory keeps, and what a tool hands back. Four mechanisms fill the window, and each one fails in its own characteristic way.

Short answer: Context engineering is the practice of deciding what a model sees on each call, and it replaced prompt engineering because the costly failures moved from wording to supply — which documents retrieval returns, what compression drops, what memory keeps, and what a tool hands back. Four mechanisms fill the window, and each one fails in its own characteristic way.

Key takeaways

  • The context window is a budget, not a container: every token spent on one mechanism is spent against every other mechanism on the same call.
  • Retrieval, compression, memory and tool return are the four ways the window gets filled, and each has a failure mode that no amount of prompt wording repairs.
  • Diagnose context as a supply chain — what enters, what is dropped, and what is allowed to persist — and the vaguest "the model got worse" report turns into a checkable claim.

context engineering: the discipline, and why it replaced prompt engineering

Prompt engineering treated the instruction as the product. You rewrote a sentence, the model behaved differently, and the loop was cheap because the rest of the input was small and stable. That assumption broke. A production call is now assembled at runtime from a system instruction, a conversation history, whatever the previous tools returned, and documents pulled from an index — and on a long agent run the instruction is a small fraction of what the model reads. The name changed because the work changed: the scarce resource is no longer how well a request is worded, it is which tokens are present at all.

Context engineering is that supply decision. For a single step it asks three questions: what the model must see, what it merely could see, and what it must not see — a retrieved document that is out of date, a tool payload carrying a credential, a memory row written by an earlier run that was never reviewed. Those are three different verbs, and most context bugs are a confusion between them: something that was merely available gets treated as required, or something required gets dropped to save tokens.

The framing that makes the discipline tractable is the budget. A window is finite, so every inclusion is also an exclusion, and the interesting engineering is not "add more context" but "decide what this step can do without". That is why the practice appears exactly when agents made context dynamic. A fixed prompt has a stable context by construction; an agent rebuilds its context at every step, and each rebuild is a chance to over-supply, under-supply, or supply the wrong thing.

None of this makes prompt engineering obsolete; it makes it a component. Wording still matters — an ambiguous tool description makes a tool the wrong choice — but wording is now judged by what it causes to be retrieved, retained and returned rather than in isolation. The practical test of a good instruction is no longer "did the model sound right" but "did it ask for the right things".

Two habits separate teams that manage this from teams that merely talk about it. The first is attribution: being able to say which mechanism put a given token in the window, because retrieval, compression, memory and tool output fail in different ways and are fixed in different places. The second is measurement before editing: a context change is a change to the input distribution, so it needs a baseline the way any other input change does. Everything below is organised around those two habits — four mechanisms, one characteristic failure each, and the smallest test that exposes it.

The four mechanisms, and the failure each one owns

The window is filled from four directions, and confusing them is why "the context was wrong" is a useless bug report. Each mechanism has one dominant failure, and the fix for a failure lives in a different place for each.

Mechanism What it puts in the window Its characteristic failure
Retrieval the documents or rows that look relevant to this step the answer existed but the query never surfaced it — recall loss, not a wrong answer
Compression a shortened form of material that is already too long the dropped sentence was the one the decision depended on
Memory facts and episodes that outlive the call a hallucination written once is retrieved as fact on every later run
Tool return the raw result of an action the model asked for the payload is larger and noisier than the answer buried inside it

Read the table as a diagnostic order rather than a shopping list. When a step produces a confidently wrong result, ask in this order: was the fact ever retrievable, did compression remove it, was a stale memory treated as current, or did a tool return the wrong slice of its own output? Each question has a different owner and a different fix, and answering them in this order is cheaper than re-prompting and hoping.

A second reason to keep the mechanisms separate is that four different parts of a stack own them. Retrieval is owned by whoever runs the index, compression by whoever assembles the request, memory by whoever owns the write path, and tool return by whoever writes a tool's output contract. A context bug reported as "the model was wrong" therefore has four possible owners, and the first job of a context investigation is to place the fault in one of the four columns before anyone edits an instruction.

The four mechanisms also trade against each other, which is the part budgets make visible. Aggressive retrieval buys recall and spends the window; compression buys window back and risks the decisive sentence; memory buys continuity and risks staleness; a talkative tool return buys completeness and drowns the signal. There is no setting that maximises all four, which is why "context engineering" is a discipline rather than a configuration.

agentic search: retrieval as a loop, not a lookup

Agentic search is what retrieval looks like once the model can issue queries instead of receiving one fixed result set. Instead of embedding the user's question once and taking the top documents, the loop forms a query, reads what came back, decides whether that is enough, and forms another — which is the difference between a search result and a search process. It buys recall on questions whose wording does not match the wording of the answer, and it spends tokens on every iteration the loop is allowed to take.

The excerpt below is the dispatch half of that machinery, cut verbatim from the product's search container. What it shows is not the model side but the plumbing a loop stands on: a resolved backend, an ordered fallback chain, and a result container that accumulates hits tagged with the engine that produced them.

# backend/smartgate/modules/search/algorithm.py — source lines 119–167 (search backend dispatch and the fallback chain)
    async def search(self):
        import time

        self.start_time = time.time()
        backend = self.settings.resolved_backend()
        self.effective_backend = backend
        try:
            if backend == "firecrawl":
                rows = await self._search_firecrawl()
                self.result_container.extend("firecrawl", rows or [])
            elif backend == "searxng":
                rows: list[dict] = []
                try:
                    rows = await self._search_searxng()
                except Exception as e:
                    logger.warning("SearXNG failed: %s", e)
                    if self.settings.searxng_fallback:
                        self.result_container.add_unresponsive_engine(
                            "searxng", str(e)
                        )
                    else:
                        raise
                if rows:
                    self.result_container.extend("searxng", rows)
                    self.effective_backend = "searxng"
                elif self.settings.searxng_fallback:
                    logger.info(
                        "SearXNG returned no rows; falling back to DuckDuckGo Lite"
                    )
                    try:
                        fb_rows = await self._search_duckduckgo_lite()
                        self.result_container.extend("duckduckgo", fb_rows or [])
                        self.effective_backend = "duckduckgo"
                    except Exception as fb_e:
                        logger.warning("DuckDuckGo Lite fallback failed: %s", fb_e)
                        self.result_container.add_unresponsive_engine(
                            "duckduckgo", str(fb_e)
                        )
                else:
                    logger.warning(
                        "SearXNG returned no rows (searxng_fallback=false)"
                    )
            else:
                rows = await self._search_duckduckgo_lite()
                self.result_container.extend("duckduckgo", rows or [])
        except Exception as e:
            logger.warning("Search backend '%s' failed: %s", backend, e)
            self.result_container.add_unresponsive_engine(backend, str(e))
        return self.result_container

Two properties in that code are worth naming, because both are failure surfaces. The fallback is ordered and narrow — Firecrawl, then SearXNG with a DuckDuckGo Lite backstop, then DuckDuckGo Lite alone — and it records which engines went unresponsive rather than silently swapping one for another, because a fallback changes the provenance of every result the loop then reasons over. And the ordinary failure is invisible in this shape: when a query simply returns nothing useful, that is not an exception, it is an empty list, and the loop must decide on its own that the query was bad rather than the world being empty. The recall problem behind that decision — query rewriting, when to stop, what a single search cannot reach — is the subject of agentic retrieval.

agentic retrieval: what changes when one search is not enough

Agentic retrieval is the discipline of deciding when another search is warranted and what query to send next. Its failure is not answering wrongly; it is answering confidently from a set that never contained the answer. A retrieval system fails in two directions, and only one of them is loud. A wrong document is usually visible — it contradicts the question. A missing document is silent, because the model reasons fluently over whatever it received and the gap only shows up as a hedged or invented detail.

That asymmetry is why the useful instrumentation for retrieval is coverage, not accuracy. For a given question you want to know whether the corpus contained a document that could have answered it at all, and whether the retrieval step surfaced it. Those are different measurements with different fixes: the first is a corpus or indexing problem, the second is a query problem. Teams that only measure "did the final answer look right" cannot tell the two apart, and end up fixing the corpus when the query was the fault.

The operational signals worth watching are cheap to collect. A retrieval step that returns nothing above a relevance threshold is a signal, not a neutral event — it should surface as a coverage gap rather than folding into an empty context. Repeated queries with near-identical wording inside one task mean the loop is not accumulating. And a final answer that cites a document which was retrieved only in an early iteration means the window is quietly re-fetching rather than remembering, which is the boundary where retrieval hands off to memory.

context compression: shrink the window without losing the decision

Compression is the mechanism that buys back window by rewriting what is already there — summarising a long document, extracting the fields a step needs, de-duplicating repeated boilerplate, or dropping low-information tokens from a verbose history. It is the mechanism most often adopted for the wrong reason: teams reach for it when they hit a hard limit, then keep the compressed text in the window after the limit stops mattering, because the summary looks tidy and nobody measures what it cost.

The failure is specific and hard to see: compression removes the sentence the decision depended on. Summaries are written for fluency, not for the next step's question, so a compressor that keeps the narrative arc can drop the one number, identifier or exception clause that a downstream tool call needs. The damage is invisible at compression time and appears three steps later as a wrong value that looks like a model error. That is why a compression change belongs in the same evaluation set as a model change — it alters the input to every step after it. The mechanics of what to drop, how to measure the loss and where a compressor belongs in the request path are worked through on context compression.

One practical rule follows from all of this. Compress material you have already decided the step does not need in full, and leave anything a decision depends on in its original form — parameters, identifiers, exact error text. A compressor that cannot tell those apart is a source of silent corruption, and the cheapest safeguard is to compress around the decision-critical fields rather than hoping the summary preserves them.

llm memory: what survives the call and what is only borrowed

Memory is the only mechanism whose subject is persistence across calls: the facts and episodes a later step is allowed to assume without retrieving them again. It exists to stop the same context being rebuilt from scratch on every turn, and it is the mechanism with the largest gap between demonstration and operation, because persistence turns a one-off error into a standing one. A wrong retrieval is wrong for one step; a wrong memory is wrong for every step that reads it afterwards.

Three distinctions make memory operable. Working context is what one task needs to finish and is short-lived by definition. Semantic memory is the durable facts — which customer prefers which format, which document is authoritative — and it is retrieved by meaning, so it needs an index and, more importantly, a policy about what may be written into it. Episodic memory is the record of what was done: the task, its steps, its outcome, which is usually the same table an observability layer already keeps. What actually gets stored, what a read is allowed to return, and what each path costs are worked through on llm memory.

The decisive operational questions are boring. What is the retention period, and does it match the promise made to a customer or an auditor? What happens when a read returns something stale — is it versioned, superseded or silently preferred? And who is allowed to write? An unsupervised write path will happily persist a hallucination and then treat it as a fact on the next run, which is why the house rule is that semantic writes are a reviewed operation rather than an autonomous one.

memory agent: the write path is the risky half

A memory agent is the component that decides what the system remembers: it reads a run's trace and writes the durable facts worth keeping. The read side gets the attention, but the write side is where the danger sits, because a write is a decision about the future taken with today's information. The failure mode is a feedback loop: an agent writes a wrong but plausible fact, a later run retrieves it, and the fact gains the authority of having been retrieved before being questioned.

Three controls keep that loop bounded, and all three are cheap. Write with provenance — a memory row should record which run produced it and from what evidence, so a later contradiction can be traced instead of hidden. Make writes reviewable rather than automatic where the fact is semantic; an episode can be written freely because it describes what happened, but a durable fact is a claim about the world and deserves the same scepticism as any other claim. And give memory the same deletion story as the audit record: retention applies to what the system remembers, not only to the logs it keeps. What an agent may decide on its own and where a human has to approve it is the subject of memory agent.

There is also a plain cost argument for keeping the write path narrow. Every stored fact is context that a future call may pay for, so memory that grows without a retention or review policy is a budget leak with a good reputation — it feels like an asset because it is retrieved, and it is retrieved because it exists.

url to markdown: the return path from a web page into context

Most of what a research step needs lives on the open web as HTML, and HTML is a terrible thing to put in a window: navigation, cookie banners, scripts and inline styling all spend tokens without carrying meaning. Converting the page to markdown is the cheapest large win in the tool-return mechanism, and it is also where the mechanism's dominant failure enters — the converted page is larger and noisier than the answer inside it, and the model has no way to know which paragraph mattered.

The pinned excerpt below is the human half of that path rather than the conversion itself: the keydown handler behind the dashboard's command surface, the keystroke that opens the dialog from which a lookup is started. It is honest to read it as the entry point — someone asks for a page — and to note that everything downstream of that keystroke, the fetch, the extraction of the main content and the markdown conversion, is where the quality is actually won or lost.

# components/dashboard/search-command.tsx — source lines 24–29 (dashboard command surface — the keydown handler)
const down = (e: KeyboardEvent) => {
      if (e.key === "k" && (e.metaKey || e.ctrlKey)) {
        e.preventDefault();
        setOpen((open) => !open);
      }
    };

The extraction half has its own failure ladder, and the order matters. A page that renders its content in the client returns a shell to anything that does not run scripts. A page that renders server-side but wraps everything in chrome returns the chrome. A converter that keeps the whole main content returns a correct but oversized document, which is a compression problem disguised as a fetch problem. And a converter that summarises aggressively returns something fluent that may not be what the page said. The mechanics — main-content selection, tables, code blocks, and what a markdown conversion should preserve — are the subject of url to markdown.

firecrawl alternative: four questions that pick the fetch layer

Fetching a page and returning it as clean markdown is now a product category, so the useful move is to compare candidates on questions that do not depend on branding. Four of them cut the list fastest.

First, what does it do with a client-rendered page — does it execute scripts, and what does it return when it cannot? Second, what is the output contract: markdown only, or markdown plus structured metadata, plus the raw HTML so you can re-extract later without a second fetch? Third, where do the fetch credentials and the per-fetch cost land — one API key in one place, or a key embedded in each service that calls it? Fourth, what happens on failure: a rate limit or a timeout should be a distinguishable outcome, because a fetch layer that returns a plausible empty document on failure is a silent recall loss dressed as success.

Run those four against a build-it-yourself option too, because the category hides a simple truth: a fetch-and-convert path is genuinely small for the common case, and a vendor earns its place at the hard edge — script-heavy pages, anti-bot surfaces, and scale. The comparison, including what a self-hosted path costs in maintenance, is worked through on firecrawl alternative.

One design rule applies regardless of the choice. Keep the fetch layer's output attributable: the caller should be able to tell which URL, which fetch, and which extraction produced the text in the window, because when a wrong fact is traced back it is usually the fetch that degraded quietly rather than the model that invented it.

Where SmartGate fits: a plan is a context budget

SmartGate is the point in our own stack where the four mechanisms meet a single enforcement surface. The gateway sees the tool call that fills the window, applies the key's rate limit and the team's token budget at the moment of the call, and writes the audit row that makes a later context investigation possible at all. That is the practical answer to "which mechanism put this token here" — the row is the record of it.

The plan table is a context budget stated in operational terms: monthly token caps of 2M, 20M, 100M and 200M-plus; MCP requests per minute per key of 120, 300, 600 and 1200; audit-log retention of 7, 30, 90 or 180 days; and 2, 10, 30 or unlimited keys per team. Read the retention column as the answer to memory's hardest question — how long a fact may be read back — and the token cap as the budget that makes compression worth doing. The pricing page is the authoritative table, and a compliance deadline should be checked against it rather than against this page.

Frequently Asked Questions

Limitations

This page is a mechanism map, not a benchmark: it names four ways a window gets filled and the failure each one owns, and it deliberately does not rank products or claim that any particular vendor implements a mechanism well. The member pages carry the mechanics, and their judgements are theirs, not necessarily ours.

The plan figures above were re-verified against the live pricing page on 2026-10-01 and are operational limits rather than a feature comparison; caps, per-key rates, retention and key counts move with the plan, so a compliance or budget decision should read the current table rather than this page.

Two of this page's eight sections quote a pinned code excerpt and six do not, because the slice matcher found no unique symbol for the other six. The consequence is stated honestly in the Method note below: there are no line-numbered claims and no quoted implementation detail for those six mechanisms — they are written from published sources and our own operational experience, not from a reading of any specific codebase.

Sources

  • Anthropic's engineering note on the discipline — Effective context engineering for AI agents, for the finite-context framing and the argument that context is a budget assembled at runtime.
  • The Prompt Engineering Guide's context-engineering chapter — promptingguide.ai/guides/context-engineering-guide, for the mechanism vocabulary used in the comparison table above.
  • IBM's topic explainer — What is context engineering?, and Gartner's Context engineering: why it's replacing prompt engineering, as the two ends of the vendor framing this page is answering.
  • LangChain's agent-context note — langchain.com/blog/context-engineering-for-agents, for retrieval, compression and memory treated as one supply problem.
  • The Model Context Protocol specification — modelcontextprotocol.io/specification, for how a tool result enters the window as context, which is the tool-return mechanism's contract.
  • Demand figures and the SERP shape quoted above are our own measurements: DataForSEO Google Ads, United States, 12-month window, measured 2026-10-01, recorded in this project's search_volume.json and research_brief.md.
  • Product behaviour and the plan table: read from the product source at the revision pinned in this project's pipeline_results.json, read-only, with the plan figures re-verified against the live /pricing page on 2026-10-01.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)
1 agentic search Search backend/smartgate/modules/search/algorithm.py 119–167 rule A L2 → slot-proof 5d88cd6aed7c
2 url to markdown down components/dashboard/search-command.tsx 24–29 rule A L2 → slot-proof 42939cf66801

Method note

This page quotes two pinned code excerpts and no others, and that split is a finding rather than a choice. The slice matcher pinned 2 of 8 sections for this page (0 abstention(s), 6 no-slice verdict(s)): rule A found a unique local symbol for agentic search and for url to markdown, and for the other six section keywords the only candidates were score-ranked and non-unique — the same generator and memory helpers surfaced under several different section keywords — so those sections are written from published sources instead. A pinned generic name would have given the page the shape of a verified article with none of the substance.

Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before publication, and the first line inside each fence records the file and the source lines it came from. The two excerpts are the search container's backend dispatch and the dashboard command surface's keydown handler; where a pin landed on the human-facing half of a mechanism rather than its core, the prose says so rather than dressing it up.

Product claims were read from the product source at the revision the slice run recorded in this project's pipeline_results.json, read-only, and the plan figures were re-verified against the live pricing page on 2026-10-01. Section keyword wording comes from this project's own paid measurement run, not from a third-party tool. No batch fingerprints, auction data or internal hosts are transcribed anywhere in this document.