SmartGate

Agentic Retrieval: The Loop That Decides When to Search

Agentic retrieval is retrieval the agent controls: it decides whether a search is needed, writes and rewrites the query, calls a tool, reads what came back, and decides whether to go again. That control flow — not the ranking model — is what separates it from one-shot retrieval, and it is why the two ways an agentic loop fails are spending too many searches and stopping too early.

Short answer: Agentic retrieval is retrieval the agent controls: it decides whether a search is needed, writes and rewrites the query, calls a tool, reads what came back, and decides whether to go again. That control flow — not the ranking model — is what separates it from one-shot retrieval, and it is why the two ways an agentic loop fails are spending too many searches and stopping too early.

Key takeaways

  • Retrieval stops being a pipeline stage and becomes control flow: four decisions — whether to search, what to ask, which tool to call, when to stop — each with its own cost and its own evidence.
  • One search call is not one lookup: the tool pages, merges, truncates and falls back behind an interface the model cannot see, so the two loops have to be budgeted separately.
  • Judge a loop on its trace rather than its answers: searches per task, distinct queries, the empty-result rate and which results the answer cites tell you whether it is over-retrieving or under-retrieving.

agentic retrieval: the loop, and what it changes about retrieval

A retrieval-augmented pipeline retrieves once. The query is the user's, the top documents are whatever the ranking produced, and generation begins with that fixed set in the window. The decision to retrieve belongs to the code, and it is made before anything is known about the answer. That shape is cheap and predictable, and it holds up while the question is single-part and worded close to the corpus.

Agentic retrieval moves that decision inside the loop. The agent asks whether it needs to look something up, writes the query itself, chooses which tool or index to send it to, reads the result, and then decides what to do with it — answer, refine the query, read one document properly, or stop. Nothing about the index changed; what changed is who owns the control flow. The retrieval step is now a decision the model can make repeatedly, with an observable cost each time it makes it.

The consequences cut both ways. The loop repairs the two questions a single pass handles badly: one whose wording does not match the answer's vocabulary, since the agent can rephrase until it does, and one with more than one part, since the agent can decompose it and check coverage part by part. It also costs more than it looks: every iteration re-sends the context already in the window, so an extra search is never the price of one call — it is that call plus everything in front of it.

The framing that keeps this manageable is the window budget — the discipline the cluster's centre page describes as context engineering. Retrieval is one of four mechanisms that fill a finite window, and the loop is the one with an iteration count, which is why the design question becomes "did this loop stop with enough evidence".

agentic search: one tool call, and the loop hiding inside it

An agent sees a search tool as a function: a string in, a list of results out. Behind that interface sits machinery the model has no view of — a resolved backend, paging, merging, truncation, a fallback chain. The excerpt below is that machinery, cut from the search container in our own stack: the paged fetch path of the SearXNG backend.

Two things in it matter for a loop. The tool has a budget of its own: the page size is fixed and the page count follows from the result count asked for, so "give me fifty results" becomes several requests rather than one. And the paging loop carries two stop rules the model never sees: a response that is not a list of results ends the loop, and a page carrying fewer than five rows is read as the end of the result set. What comes back is then truncated to the requested count.

# backend/smartgate/modules/search/algorithm.py — source lines 236–278 (search container — paged SearXNG fetch and its two stop rules)
        per_page = 20
        want = self.search_query.max_results
        pages_to_fetch = max(1, (want + per_page - 1) // per_page)
        start_page = self.search_query.pageno

        merged: list[dict] = []
        async with httpx.AsyncClient(timeout=30.0) as client:
            for offset in range(pages_to_fetch):
                page = start_page + offset
                params: dict[str, Any] = {
                    "q": self.search_query.query,
                    "language": self.search_query.lang,
                    "pageno": page,
                    "format": "json",
                }
                if engines_param:
                    params["engines"] = engines_param
                if categories_param:
                    params["categories"] = categories_param

                resp = await client.get(
                    final_url,
                    params=params,
                    headers={"User-Agent": UA, "Accept": "application/json"},
                )
                resp.raise_for_status()
                data = resp.json()
                chunk = data.get("results") if isinstance(data, dict) else None
                if not isinstance(chunk, list):
                    break
                for a in chunk:
                    merged.append(
                        {
                            "title": (a.get("title") or "").strip(),
                            "url": (a.get("url") or "").strip(),
                            "content": (a.get("content") or "").strip(),
                            "engine": "searxng",
                        }
                    )
                if len(chunk) < 5:
                    break

        return merged[:want]

The consequence is easy to miss: there are two loops inside one "search". The inner one belongs to the tool, stops on its own rules, and can return a short list without raising anything; the outer one is the agent's, and its only evidence is the shape of the result set. That is why the query passes through untouched here: the container ships the words it was given, and rewriting them is the agent's job. An empty list is data, not an exception; what the caller does with it is the loop.

The loop in four moves

It is worth writing the loop down as four decisions rather than as a prompt, because each move takes a different input, produces a different artefact, and fails differently.

Move The decision What it reads How it goes wrong
Whether to search is this answerable from the task, memory or the model itself? the question's specificity, a memory hit, a freshness requirement searching reflexively for everything, or never searching at all
What to ask which query, against which index, with which filters the user's wording, earlier results, the part that is still missing restating the question verbatim, or sending one query for a multi-part question
Which tool to call which search or fetch tool, with what budget and timeout the tool's own contract, its cost, whether it is answering at all unbounded fan-out, or a fetch that returns a plausible empty page
When to stop is the evidence sufficient, contradictory, or absent? coverage of every part of the question accepting a fluent summary, or stopping because the model said it was done

Read as a state machine, the loop is small enough to draw: each move writes one line into the run's trace, and the next reads that trace rather than the conversation. That is the difference between a loop and a retry cycle — a retry re-runs the same request and hopes for a different outcome, while an explicit loop can explain afterwards why it stopped where it did.

Two notes follow. The cheapest move to improve is usually query wording, because it decides what the first call can return; the most often broken is the fourth, since a loop with no sufficiency test cannot tell a complete answer from a confident one. And whether to search at all, and when to stop, are the two moves ordinary code can gate without losing anything the model contributes.

Stopping conditions: what can actually end the loop

An agent's judgement that it has enough is a useful signal and useless as a guard: the step that wants another search is the step that would have to enforce the limit. Three conditions can be enforced from outside the prompt, and a loop that ships should carry all three.

A budget. Count searches per task, not only tokens: a search count is the number the loop can read before it acts, and the one that maps onto cost. A small explicit cap — three to five searches for a factual task — is what makes cost predictable, and it should be raised from evidence: correct answers at the cap mean extra searches buy latency, while wrong answers at the cap usually mean query coverage, not the cap.

A sufficiency test against the question. The move that ends most loops badly is checking the answer's fluency rather than whether every part of the question has support. Write the parts down — two for a comparison, three or four for a research brief — and require a source for each. It is the only condition of the three that speaks to correctness, and it reads the trace the loop already produces.

A no-progress test. Two failures are indistinguishable from progress unless you look for them: a query near-identical to one already sent, and a result set that came back empty twice. Treat an empty result as information and change the query rather than repeating it, then stop after the second empty return with the gap stated. A loop that ends by naming the fact it could not find beats one that keeps searching, and beats one that fills the gap itself.

Two smaller habits make the conditions workable: a wall-clock or step ceiling no component can extend, and enforcement outside the prompt — budget and stop logic belong to the code or the gateway, because a model that can extend its own budget eventually will.

Two failure modes: over-retrieval and under-retrieval

Loops do not fail symmetrically: the two directions need different numbers to detect and different fixes to correct, so naming which one you have is most of the diagnosis.

Over-retrieval is the loop spending calls to avoid thinking. Its symptoms show up in aggregate: searches per task rising over time, tokens per answer rising faster than quality, the same fact fetched on three consecutive iterations, and — the one that hurts most — a window holding near-identical documents that disagree in detail, leaving the model to choose on confidence rather than evidence. The causes are structural: no budget, no de-duplication of queries, and a treatment of "one more search" as cheaper than compressing what is already there. The fixes are a cap, a query ledger, and a merge step that reconciles contradictions first.

Under-retrieval is the loop stopping on the first plausible set. It is quieter, because the answer reads well. Its symptoms: one search per task however compound the question is, an answer that cites nothing, an empty result treated as proof that no answer exists, and sub-questions dropped between iterations. One cause does not look like retrieval at all: the agent that answers from its own parametric memory without searching, because the question felt familiar. That is the decide move failing, and it presents as a hallucination.

The measurements that separate them are cheap: searches per task and the duplicate-query rate diagnose over-retrieval, while coverage — the fraction of the question's parts backed by something the loop actually read — diagnoses under-retrieval, provided the question was decomposed first. The fixes route differently too: budgets, de-duplication and compression for the first; decomposition, a coverage check and a rule that no claim arrives unsupported for the second. Where the missing support is a stale fact rather than absent evidence, the failure has left the loop and moved into memory.

context compression: the alternative the loop should price first

When the window is the binding constraint, the loop has two options that look different and compete for the same budget: fetch less, or shrink what it already holds. Compression is the second, and it is the option teams reach for at a limit and then keep doing after the limit stops mattering.

Inside a loop, the useful distinction is between compressing history the loop has decided not to re-read in full, and compressing a freshly returned document whose details the next call may need. The first is safe and buys window back on every later iteration. The second is where compression quietly changes the answer: a summary written for fluency can drop the identifier, the number or the exception clause that the next step depends on, and the damage appears two moves later as a wrong value that looks like a model error. A truncated tool result is the same decision taken by whoever wrote the tool. What to drop, and how to measure what the drop cost, is worked through on context compression.

llm memory: the loop's second retrieval path

The decide move does not have to resolve to a search. For a fact the system has already learned, the right answer is a memory read — cheaper than a search, and repeatable without a tool call. Treating that read as just another form of retrieval is what keeps the loop honest, because memory has the same shape as an index and a different failure: it can be stale rather than missing, scoped to the wrong tenant or session, or written once from a single bad inference and retrieved as a fact ever after.

The loop's job is to make those three distinguishable: a memory hit should arrive with an age and a source, the trace should record whether a fact came from memory or a search, and fresh evidence should be allowed to overrule memory. The mechanics — what gets stored, what a read may return, and what each path costs — belong to llm memory.

memory agent: what the loop is allowed to write back

Every loop that stops also decides, somewhere, what to keep. The cheapest safe split is to write episodes freely and semantic facts carefully: what was searched, which tool ran, what came back, and how the loop ended are all records of what happened, and they are exactly what the next iteration needs to avoid repeating the last one. A durable fact is a different object — a claim about the world — and it needs provenance, and usually review, because an unattended write turns one plausible inference into a standing assumption that every later run trusts.

Keep the write path small for a second, plainer reason. Everything remembered is context a future call may pay for, so a memory that grows without a retention or review policy is a budget leak with a good reputation. What an agent may decide on its own, and where a human has to approve it, is the subject of memory agent.

firecrawl alternative: the tool choice is a loop decision

Because the model cannot see inside the tool, the tool's failure modes become the loop's. That makes the vendor question a design question rather than a procurement detail, and four questions settle most of it.

Does it paginate or truncate, and can the caller tell which happened? Paging that stops early and truncation that is invisible both look like a short result set from the outside, and a loop cannot tell "the web has little on this" from "the tool stopped at ten rows". Does the output carry provenance — which URL, which fetch, which extraction produced the text — because a wrong fact is traced back through the tool far more often than through the model? Does an error look different from an empty result, or does a rate limit arrive as a plausible empty page? And where does the key live, with what per-call cost, since a loop multiplies whatever one call costs?

Run the same four against a self-hosted path before buying anything, because the common case is small: fetch, extract, convert, return. A vendor earns its place at the hard edge — script-rendered pages, anti-bot surfaces, and scale — and the trade-off, with the maintenance cost of self-hosting, is set out on firecrawl alternative.

url to markdown: what the loop gets back into the window

Search returns pointers; the loop usually needs text. Page-to-markdown is where a result list becomes something a model can reason over, and where much of the tool-return failure enters: HTML spends tokens on navigation, cookie banners, scripts and styling that carry no meaning for the question.

The pinned excerpt below is the entry point rather than the conversion: the keydown handler behind the dashboard's command surface, the keystroke that opens the dialog a lookup starts from. Read it as the human half of the path — someone asks for a page — and note that everything downstream is where quality is won or lost.

# components/dashboard/search-command.tsx — source lines 24–29 (dashboard command surface — the keydown handler)
const down = (e: KeyboardEvent) => {
      if (e.key === "k" && (e.metaKey || e.ctrlKey)) {
        e.preventDefault();
        setOpen((open) => !open);
      }
    };

The failure ladder is worth knowing in order: each rung fixes a different problem. A page that builds its content in the browser returns a shell to anything that does not run scripts. A server-rendered page wrapped in chrome returns the chrome. A converter that keeps the whole main content returns a correct but oversized document — a window problem wearing a fetch costume. And one that summarises aggressively returns fluent prose that may not be what the page said: the rung that produces wrong answers rather than wasteful ones. The mechanics are covered on url to markdown.

agent memory systems: the ledger that turns retries into a loop

An agentic loop needs somewhere to record what it has already asked and learned. Without that record, iteration three is iteration one with a higher token bill, and the stopping conditions have nothing to test against: no sufficiency check runs without knowing what is still missing, and no no-progress test fires without knowing which queries were already sent.

The useful ledger is small and boring. Per iteration: the query as sent, the tool and backend that answered it, how many results came back, which URLs the loop actually read, and the decision it took afterwards. Two payoffs follow immediately: the loop stops re-asking, and the trace becomes readable to someone who was not in the conversation. A third follows from retention, because this ledger is operational history: how long it is kept is the same question you answer for logs, and it sets how far back a run can look before it must search again.

That is the sense in which the surrounding mechanisms make up a system rather than a stack of features: retrieval puts evidence in the window, the ledger stops the loop re-fetching it, memory carries it across runs, and compression keeps all of it inside the budget. It is also why the two failure modes above are so often a ledger problem in disguise — an agent that keeps searching is usually an agent that cannot see what it already found.

Where SmartGate fits: the loop's enforcement boundary

Every call a loop makes can cross one place in our own stack, where two of the four moves stop being prompt-level intentions. The gateway sees the tool call when it happens, applies the key's per-minute request limit and the team's token budget at that moment, and writes the audit row a loop needs anyway — which query, which tool, how many results, how many tokens, how long it took, and how it ended. The ledger above is therefore the default output of that layer rather than a second system to build, and the same row is what a later investigation reads to find which iteration supplied a wrong fact.

The plan table is where those limits are stated in operational terms: monthly token caps of 2M, 20M, 100M and 200M-plus; MCP requests per minute per key of 120, 300, 600 and 1200; audit-log retention of 7, 30, 90 or 180 days; and 2, 10, 30 or unlimited keys per team. Read the retention column as the outer bound on how long a loop may consult its own history, and the per-key rate as the ceiling on parallel fan-out. The pricing page is the authoritative table.

Frequently Asked Questions

Limitations

This is a design page for a control-flow pattern, not a benchmark. It ranks no vendor and quotes no measurement of retrieval quality, because the numbers that matter here are properties of your own traces: searches per task, duplicate-query rate, empty-result handling and coverage of the question's parts. Two implementations that both call themselves agentic retrieval can differ more than either differs from a fixed pipeline, so read the four moves as a checklist rather than as an architecture.

Two of this page's eight sections quote a pinned code excerpt and six do not, because the slice matcher found a unique local symbol for only two keywords. One of the two excerpts is the human-facing entry point to a lookup rather than retrieval itself, and the Method note says which and why. For the six unpinned sections there are no line-numbered claims and no quoted implementation detail; those statements come from published sources and from operating this stack.

The plan figures above were re-verified against the live pricing page on 2026-10-01 and are operational limits rather than a feature comparison: caps, per-key rates, retention and key counts move with the plan, so a compliance decision should read the current table. Demand figures are our own single-market measurement and describe the vocabulary this page answers, not the traffic an implementation can expect.

Sources

  • Anthropic's engineering note on agent design — Building effective agents, for the loop-versus-workflow split and when iteration earns its cost at all.
  • Anthropic's account of a production research loop — How we built our multi-agent research system, for fan-out search budgets, the token cost of parallel retrieval, and stopping rules in practice.
  • ReAct — arxiv.org/abs/2210.03629, for interleaving reasoning with tool calls, which is the loop shape this page is written around.
  • Self-RAG — arxiv.org/abs/2310.11511, for retrieval decisions and self-critique produced by the model rather than fixed by a pipeline.
  • IRCoT — arxiv.org/abs/2212.10509, for interleaved retrieval and reasoning on multi-step questions, the case where one query is provably not enough.
  • The Model Context Protocol specification — modelcontextprotocol.io/specification, for how a tool result is returned to a caller.
  • Firecrawl's documentation — docs.firecrawl.dev, as the shape of a fetch-and-convert tool contract.
  • Demand figures and the SERP shape quoted above are our own measurements: DataForSEO Google Ads, United States, 12-month window, measured 2026-10-01, recorded in this project's search_volume.json.
  • Product behaviour and the plan table: read from the product source at the revision pinned in this project's pipeline_results.json, read-only, with the plan figures re-verified against the live /pricing page on 2026-10-01.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)
1 agentic search Search backend/smartgate/modules/search/algorithm.py 236–278 rule A L2 → slot-proof 5d88cd6aed7c
2 url to markdown down components/dashboard/search-command.tsx 24–29 rule A L2 → slot-proof 42939cf66801

Method note

This page quotes two pinned code excerpts and no others, and that split is a finding rather than a choice. The slice matcher pinned 2 of 8 sections (0 abstention(s), 6 no-slice verdict(s)): rule A found a unique local symbol for agentic search and for url to markdown, and for the other six section keywords the only candidates were score-ranked and non-unique — the same compressor and memory helpers surfaced under several different section keywords — so those sections are written from published sources instead. A pinned generic name would have given the page the shape of a verified article with none of the substance.

Every fenced block above was cut out of the slice body and re-asserted against it byte-for-byte before publication, and the first line inside each fence records the file and the source lines it came from. The two excerpts are the search container's paged fetch path, with its own two stop rules, and the dashboard command surface's keydown handler. The second is the human entry point to a lookup rather than the retrieval itself, and it is presented as such above.

Product claims were read from the product source at the revision the slice run recorded in this project's pipeline_results.json, read-only, and the plan figures were re-verified against the live pricing page on 2026-10-01. Section keyword wording comes from this project's own paid measurement run, not from a third-party tool. No batch fingerprints, auction data or internal hosts are transcribed anywhere in this document.