Deep Research Agent Architecture: Planner, Searcher, Verifier
A deep research agent is one question run through three roles — a planner that decides what to look for next, a searcher that fetches and normalises evidence, and a verifier that checks every drafted claim against a retrieved span before it survives. The roles are the architecture; the stop condition, the budget and the verification pass are the engineering.
Short answer: A deep research agent is one question run through three roles — a planner that decides what to look for next, a searcher that fetches and normalises evidence, and a verifier that checks every drafted claim against a retrieved span before it survives. The roles are the architecture; the stop condition, the budget and the verification pass are the engineering. This page is about how to wire the three together and how to keep the loop from collapsing, drifting or confirming itself.
Key takeaways
- A deep research agent is not one long prompt: it is a planner, a searcher and a verifier with typed handoffs, and each role needs a different failure policy.
- The loop has to terminate on counters — steps, tokens, wall clock and new distinct sources — never on the model's own claim that it has finished.
- Retrieval collapse, citation drift and self-confirmation are the three failure modes that survive good prompt engineering, and each one has a mechanical guard rather than a wording fix.
- Budget belongs at the run level and the role level at the same time; a per-call timeout without a run-level deadline is how a research job keeps billing until the morning.
- Write the claim ledger and the stop counters before you write the planner — those two artefacts are what make the rest of the agent testable instead of impressive.
deep research agent: the planner, the searcher and the verifier
A deep research agent answers a question it was not handed the sources for. That property is what separates it from a retrieval-augmented pipeline and from a chat assistant: there is no index to read, so the system has to decide what to search for, go and search, and then decide whether what came back is enough. The design that has survived production is a triangle of three roles, each with a narrow job and a typed output.
The planner owns the question. It turns one sentence into a short list of sub-questions, an order to ask them in and an acceptance condition for each — what evidence would make that sub-question answered. The plan is data, not prose: a list of typed items a log can print and a test can assert. The searcher owns coverage. It takes one sub-question and returns results normalised into a single shape — a title, a url and a text span — so everything downstream reads one contract instead of five. It decides what to fetch and how many times; it does not decide which sentences the answer will use. The verifier owns the claim ledger. Every sentence the writer wants to emit becomes a claim, the verifier checks it against the span that supposedly supports it, and it records pass, fail or unsupported. The verifier never writes prose.
The interface between the roles is where these systems are won or lost. If the searcher can silently reinterpret the question, or the verifier can rewrite a claim instead of grading it, the three roles collapse back into one model with a longer prompt — and every failure mode below comes back with it. The busy tool layer this triangle sits on is described across the cluster, and the single most useful first stop is the AI research tool landscape, which maps what each class of tool actually decides before you buy one. What the loop itself does step by step, and how the roles share one context window, is the mechanism page's job — see how the deep research AI loop works. This page stays on construction.
ai research agent: what each role owns and what it may not touch
Three roles are easy to name and easy to blur, so give each one an explicit contract: what it receives, what it returns, what it is allowed to change, and what happens when it fails.
| Role | Receives | Returns | May not | On failure |
|---|---|---|---|---|
| Planner | the question, the budget, the source ledger so far | a ranked list of sub-questions, each with an acceptance condition | fetch, or mark its own plan complete | re-plan once, then hand back what the plan produced |
| Searcher | one sub-question | normalised hits: title, url, span, engine | reinterpret the question, or summarise for the writer | record the engine as unresponsive and return the partial set |
| Verifier | a claim and the span that supports it | pass, fail or unsupported, with a reason | rewrite the claim or fetch a rescue source | mark the claim unsupported so the writer drops it |
The handoff rules matter more than the role names. First, the plan is the only object the planner may change; a searcher that quietly narrows or widens the question is where retrieval collapse starts. Second, partial results are normal: the searcher returns what it got and names what failed, rather than raising and killing the run. Third, the verifier grades and never repairs — a verifier allowed to rewrite a claim will rewrite it toward agreement, and at that point you have two writers and no check. Fourth, exactly one component merges: the writer consumes the ledger and may only emit sentences whose claims are marked pass.
That division is also what makes an ai research agent testable rather than merely impressive. A plan is a list, a hit is a record, a claim is a row with a status — all three can be asserted in a unit test without a model in the loop. When the design does reach for a model again, the discipline is the same one an evaluation harness uses: a case has an input and something checkable in the output. A claim that must cite a span is cheap to check mechanically, which is why the ledger is worth building before the planner. When the source of a claim is a paper, the check the verifier performs is the same one a reader performs when reading an AI research paper — does the sentence say what the span says, in the paper's own terms.
autonomous research agent: stopping the loop on purpose
The word "autonomous" is where budgets go to die, because the most natural stop condition — the model saying it is done — is the one condition you cannot trust. A model asked whether it has finished has no incentive to say no, and the same pass that produced the answer is the pass judging it. So the termination contract is written in counters, and the model's opinion is an input rather than the rule.
Four counters cover almost every real run, and they are cheap to hold in one state object. A step cap bounds how many search and read actions the run may take. A token budget bounds what the whole run may spend, checked before each call rather than reconciled after it. A wall clock bounds the run's age end to end. And a novelty floor bounds the marginal value of continuing: the run tracks new distinct sources added per step and stops when two consecutive steps add fewer than one. That last counter is the one people forget, and it is the one that stops a loop which is technically making progress while adding nothing.
Two rules make the counters work. Termination is decided in deterministic code between steps, never inside a prompt — a stop condition the model can argue with is advice, not a condition. And every counter is recorded on the run, so "why did this job take forty minutes" has an answer that is not a guess. This is the same shape a scheduler uses everywhere else: a bounded loop with an explicit exit, except that here the work inside the loop is non-deterministic and therefore the exit has to be more explicit than usual, not less. When a run stops on a counter rather than on "answered", say so in the report: an honest truncated answer with its stop reason beats a complete-looking one nobody can audit.
multi agent research: splitting the budget and the clock
Once the three roles are real, their cost is no longer one number, so the budget has to be layered. There is a run budget — the ceiling the whole task must not cross — and a role budget, the slice each role may draw from it. The layering is not bureaucracy: a planner that spends the run's budget has nothing left for the searcher it planned for, which is the most common way a research job produces a beautiful plan and no evidence.
Time needs the same treatment, and this is where a naive implementation fails quietly. A per-call timeout protects one request; it does nothing about a run that makes two hundred healthy requests. So there is a per-call timeout (how long one fetch may take), a per-role budget (how long a role may spend), and a run deadline (when the whole thing stops regardless). The deadline is a stop condition of the previous section, not an error path — the difference between "we ran out of time and here is what we have" and "the process died" is the difference between a usable partial answer and a job you must pay for twice.
Degradation is a policy, and it should be chosen rather than defaulted. When a search backend is slow or returns nothing, fail-open means record the engine as unresponsive, keep the partial evidence and continue; fail-closed means stop the run because coverage cannot be assured. Most research runs want fail-open with a loud ledger entry, because a missing slice of coverage is a caveat, not a failure — but a run whose whole point is a defensible number wants fail-closed, because a partial evidence set presented as complete is worse than no answer. What the split costs in tokens, and where the fan-out multiplier lands, is the operational half of multi agent research — the point here is only that both the money and the clock get a run-level ceiling before any role-level one.
deep research ai: the three ways a run goes wrong
A deep research AI run fails in three characteristic ways that survive good prompting, because they are properties of the loop rather than of any one instruction. Knowing the symptom is what lets you pick the guard.
Retrieval collapse is the loop converging on its own first hypothesis. The planner writes a query, the searcher returns a cluster of sources that all say the same thing, the next query is phrased from those sources, and the run narrows instead of broadening. The symptom is high activity with low diversity: many fetches, few distinct domains, and every new source echoing the last. It happens because each step is conditioned on the previous context, and a context full of one view asks questions that view would ask.
Citation drift is a claim outliving its source. The span is real and the url resolves, but the sentence has travelled: a caveat was dropped, a correlation became a cause, a specific number became "roughly". The symptom is a ledger where every claim passes but the citations, read end to end, do not quite say what the report says. It is the quietest of the three, because every individual link looks healthy; only reading claim and span together exposes it.
Self-confirmation is the verifier agreeing with the writer. It happens when both share a model, a context or a prompt, so the checker inherits the generator's assumptions and returns a high pass rate on exactly the claims that should fail. The symptom is a suspiciously clean ledger — a hundred per cent pass is not a sign of quality, it is a sign that the check is not independent. A related trap is that the run's own summary becomes a source: once a claim is in the context, the next step can "retrieve" it from the summary rather than from the world.
ai research tool: the guards that hold the loop in engineering terms
Each failure mode has a mechanical guard; none of them is a wording fix, and all of them are cheap to build once you know which one you are building.
| Failure mode | Guard | What it changes |
|---|---|---|
| Retrieval collapse | a query-diversity floor: every plan must carry a minimum number of distinct intents, and the ledger tracks distinct domains, not fetches | the planner, not the searcher |
| Citation drift | source-anchored claims: a claim is a row holding the url plus the quoted span, and it cannot be emitted without one | the writer's output contract |
| Self-confirmation | independent verification: a checker without the generator's context, ideally a different model, plus a claim-level pass rate that is monitored rather than trusted | the verifier's isolation |
| All three | dedup and compression before context grows, so a narrowing loop cannot crowd out its own contradicting evidence | the context layer |
| Runaway cost | the counters and the layered budget above, enforced in code between steps | the loop controller |
Two of these deserve a note. Deduplication is usually filed under cost, but it is also a correctness guard: if the same passage arrives five times through five urls, a context that keeps all five amplifies one source into apparent consensus, which is retrieval collapse wearing a disguise. And a monitored pass rate is the honest version of verification: the number of failed claims per run is a signal to watch, and a run that never fails a claim is a run to inspect, not to trust. The primitive set that these guards are usually assembled from — a fetch that normalises a page, an aggregate search, a dedup pass, a context gate, a memory store, a budget guard and a pipeline that orders them — is the tooling layer; the failure modes above are the questions to ask each class of tool before adopting one.
ai researcher: what the agent still hands back to a human
A well-built agent does not close the loop on everything, and pretending otherwise is how a research tool loses trust on its first hard question. There is a class of work the loop should hand back, and it is worth naming so the handback is a feature rather than an apology.
The agent reads what is publicly reachable. Sources behind a sign-in, a paywall, a database licence or an internal drive are invisible to it, which for a literature-heavy question can be most of the relevant material. It reads what is indexed now, so a document published this morning may not exist for it, and a retracted paper may still appear. It reasons over spans it retrieved, so it is good at "what does this source say" and weak at "which of these two reports is right" when the answer needs domain judgement or a causal argument no span states. And its recall on the long tail is uneven: the canonical survey and the famous blog post come back every time, while the one obscure paper that actually answers the question may not surface at all.
Those limits translate into three handback rules. A decision with legal, financial or clinical weight keeps a human at the merge point, because the agent can assemble the evidence but cannot own the call. A question with contradictory sources goes to a person by default, since picking a side is judgement and the loop has no ground truth to pick with. And a run that returns very few distinct sources should say so at the top of its report, so the reader knows the confidence before they read the conclusion. The role that receives this handback is the analyst, and what that role looks like day to day — the reading, the checking, the judgement the loop cannot supply — is described on the AI researcher role. Where the work is volume rather than judgement, a structured AI literature review is exactly the part of the job the agent should take over completely, with a person reviewing the extraction rather than performing it.
Which engines a research run actually calls, read from our own module
A research agent's answer is only as multi-source as its search step. Here is ours, read on 2026-10-07
from backend/smartgate/modules/search/algorithm.py and modules/search/__init__.py.
- The provider is a setting with four values:
auto(the default), or a forcedfirecrawl,searxngorduckduckgo. A forced provider is a useful debugging lever and a quiet single point of failure in production — if you pin one engine, you have pinned one engine's rate limits and outages too. - A self-hosted SearXNG instance is a first-class option, with its own URL, optional engine and category lists, and a fallback switch that is on by default. The fallback is the difference between "our instance is down, the agent stops" and "our instance is down, the agent keeps going with lower quality".
- Results keep their provenance. Each hit carries the engine or engines that produced it and the positions it appeared in, and the module normalises that set before the result leaves the tool boundary — because a raw set is not JSON, and a serialisation failure at the edge is a silent loss of the very field that tells you where a claim came from.
- The tool is a module with declared dependencies and a version. The search surface is versioned like the rest of the gateway, so "which search behaviour produced this run" is answerable after the fact rather than reconstructed from memory.
Where SmartGate fits
SmartGate is the intelligence layer this kind of construction runs on, and the honest reason to put a gateway there rather than call search and fetch directly is the same list as above: the record, the budget and the guard all need one place to live.
| Shape | Who decides the stop | Cost profile | What it gives up |
|---|---|---|---|
| A single long prompt | the model, implicitly | one call, unpredictable under retry | verification, and any audit of how the answer was reached |
| A self-built planner / searcher / verifier | your loop controller, on counters | per-step, bounded by your budget code | the maintenance of the loop itself |
| A managed deep research API | the vendor's loop | per-task, opaque granularity | control of the guards, and the raw evidence trail |
The gateway closes the gap in the middle row. Every tool call — search, fetch, dedup, context, memory —
is written as an audit row with the caller, the route, the token count and the outcome, so the claim
ledger and the step counters read the same record the budget reads. smart_budget_guard gives the run
the ceiling and the role its slice, smart_context_gate and smart_dedup keep the context from
crowding out contradicting evidence, smart_memory holds what a later run is allowed to assume, and
smart_pipe orders the research and remember templates so the loop controller does not have to
reimplement them. Plans start at a few dollars a month, with monthly token caps of 2M, 20M, 100M and
200M+, requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180
days, and a team holding 2, 10, 30 or unlimited keys; the billing model shares only in what the
platform saves, and the pricing page is the authoritative table —
read the numbers there rather than this sentence.
To start, add the gateway as an MCP server and point your existing search and fetch calls at it: the
loop you already wrote keeps its shape and gains a record and a ceiling. When you want the
research-and-remember path pre-ordered, smart_pipe runs it in one call. You can
start free without a billing step, and if the
pattern you are building is a service rather than an application, the same primitives are what a
deep research API exposes once the loop is wrapped behind an endpoint.
Frequently Asked Questions
Is a deep research agent just a long prompt with a search tool?
No. A long prompt with a search tool has one role that decides, searches and answers, so the same generation that writes a claim also judges it. A deep research agent splits those into a planner, a searcher and a verifier with separate outputs, which is what makes the claim checkable and the stop condition enforceable in code rather than in wording.
How many steps should a research run be allowed?
Enough to cover the plan and no more. In practice a step cap in the low tens, a run token budget, a wall-clock deadline and a novelty floor that stops the loop when new distinct sources dry up. The exact numbers depend on the question, but the rule does not: stop on counters, not on the model's opinion that it has finished.
Why does a clean verification pass worry you?
Because a verifier sharing the writer's model or context inherits the writer's assumptions, so it tends to pass the very claims that should fail. One hundred percent pass is a symptom of a check that is not independent, not a sign of quality. Isolation — a different model, and never the generator's context — is what makes the pass meaningful.
Do we need three separate models for the three roles?
No, you need three separate contracts and, for the verifier, isolation. The planner and searcher can be the same model with different instructions; the verifier is the one role where sharing the writer's context costs correctness, so give it a fresh context and ideally a different model.
What should the agent refuse to do on its own?
Own a decision with legal, financial or clinical weight, resolve directly contradictory sources, and claim coverage it does not have. It can assemble the evidence for all three; the merge point stays with a person, and a run that found very few sources should lead with that fact rather than bury it.
Limitations
This page describes an agent architecture, not a specific product, and it does not rank any tool: the right split of planner, searcher and verifier depends on your sources, your budget model and how much of the record you are allowed to keep. The failure modes and guards are the ones that recurred often enough to be worth naming; they are not exhaustive, and a domain with unusual retrieval constraints will have its own.
The architecture cannot fix a thin source base. If the answer depends on material behind a sign-in or on a document published minutes ago, no amount of loop engineering will retrieve it, and the honest outcome is a clearly-scoped partial answer. It also cannot make a weak verifier strong: an isolated check tells you when a claim fails, but it does not certify that a passing claim is true, only that the span supports it as written.
Finally, the page carries no code excerpt, and the Method note below records why. What follows from that honestly: no line-numbered claims, no quoted implementation detail, and no assertion about the internal structure of any named project.
Sources
-
Anthropic's engineering note on agent patterns — Building effective agents, for the workflow-versus-agent split and the reason a fixed pipeline beats an open loop when the sequence is known.
-
The ReAct paper — arXiv:2210.03629, for interleaving reasoning and tool use, which is the shape the planner-searcher loop formalises.
-
Chain-of-Verification — arXiv:2309.11495, for independent verification of generated claims and why the checking step must not share the draft's context.
-
Self-Refine — arXiv:2303.17651, as the self-critique baseline the isolated verifier is deliberately not: same-model feedback helps fluency and does not close the self-confirmation gap.
-
Lost in the Middle — arXiv:2307.03172, for why evidence buried in a long context is used less, which is the mechanism behind dedup and context compression as a correctness guard.
-
OpenAI's deep research announcement — openai.com/index/introducing-deep-research, and the Model Context Protocol specification at modelcontextprotocol.io, as the wire contract the tool layer in the comparison table speaks.
-
Demand figures on this page are this project's own paid measurements, recorded in its
search_volume.jsonandresearch_brief.md; product behaviour was read read-only from the product source at the revision this project's slice run recorded. -
The section "Which engines a research run actually calls, read from our own module" is our own implementation, read on 2026-10-07 from
backend/smartgate/modules/search/algorithm.pyandmodules/search/__init__.py(origin/main). It states only what those files state.
Method note
This page carries no code excerpt, and the reason is a measurement rather than a preference. The slice matcher pinned 7 of 7 sections for this project (0 abstention, 0 miss), but every pinned section resolved to the same general-purpose asset — a web-search container class whose name collides with ordinary search vocabulary across a product codebase. One generic symbol standing behind seven different section topics is not section-level evidence: dressing it up as seven excerpts would have given the page the shape of a verified article with none of the substance. So the house rule for an unpinned section applies here — write from sources, invent nothing — and the page is deliberately code-free.
The product behaviour above (the primitive set and the plan limits) was read read-only from the product source at the revision this project's slice run recorded, and the plan figures were re-verified against the live pricing page on 2026-10-02. The demand figures come from this project's own paid keyword measurement, not from a third party. No code, batch fingerprints, auction data or internal hosts appear on the page, so there is nothing here that has to be asserted verbatim.