SmartGateSmartGate

Deep Research AI: The Stages Inside the Loop

Deep research AI is a staged system, not a longer answer: it plans before it searches, retrieves in rounds, reads each page, keeps a working set, drafts from the passages it kept, checks every citation against what it actually read, and then decides whether to retrieve again. The loop is slower and dearer than one-shot long-context QA because every stage adds calls.

Short answer: Deep research AI is a staged system, not a longer answer: it plans before it searches, retrieves in rounds, reads each page, keeps a working set, drafts from the passages it kept, checks every citation against what it actually read, and then decides whether to retrieve again. The loop is slower and dearer than one-shot long-context QA because every stage adds calls. It earns that cost when the answer has to be defended, not merely fluent.

Key takeaways

  • A deep research run is a pipeline of six stages — plan, retrieve, read, synthesise, verify, re-retrieve — and each stage is a separate place the run can be stopped, billed or audited.
  • It differs from one-shot long-context QA in three measurable ways: more retrieval calls, a working set that survives between rounds, and citations that are checked rather than produced.
  • The bill is not one bigger prompt. It is planning tokens plus retrieval calls plus page reads plus verification passes, and it grows with the number of rounds, not the length of the question.
  • Held against classic RAG, the difference is who chooses the next query: a RAG pipeline fixes it before the model runs; a deep research loop chooses it after reading the last round.
  • Before you run one, write down the ceiling first — maximum rounds, maximum pages read, maximum spend — because a loop with a ceiling degrades into a smaller report, and a loop without one degrades into an invoice.

The category is easy to misread, because the visible half is a chat box and the invisible half is a pipeline. This page stays on the invisible half: it takes the loop apart stage by stage, shows where our own retrieval layer sits inside it, and states plainly when the machinery is buying something a single search call could not.

deep research ai: the stages between a question and a cited report

A deep research AI system is best read as six stages, because each one is separately observable and separately priced. Plan: the question is restated as a handful of sub-questions and a strategy is chosen. Retrieve: searches are issued for those sub-questions. Read: the pages behind the top results are fetched and the passages that matter are extracted. Synthesise: a draft is written from the passages that were kept, not from the whole web. Verify: the draft's own citations are walked and each claim is checked against the text that was actually read. Re-retrieve: the gaps in the draft become new queries, and another round runs until a stop condition is met.

Two properties separate this from a chat answer. First, the stages are sequential and stateful: what the read stage keeps decides what the synthesis stage can say, and what the verification stage finds decides whether another round is worth running. Second, the artefact is a report with a visible evidence trail, so a reader can disagree with one sentence and follow the citation back to a page instead of re-arguing a whole answer.

Stage The decision it owns What it costs
Plan What the sub-questions are, and how deep to go One model call, before any evidence exists
Retrieve Which queries run, and how many results each returns One search call per query, per round
Read Which pages enter the working set, and what is kept from them One fetch per page, plus the tokens to hold the extract
Synthesise What the draft claims, and which passages are cited One long generation over the working set
Verify Whether each citation supports its sentence A pass over the draft and its sources
Re-retrieve Whether the gaps justify another round The whole cost above, multiplied

The retrieval stage is where the loop meets infrastructure, and the sensible way to build it is a fallback chain rather than a single API. The excerpt below is the entry point of our own retrieval container: it resolves a backend from settings, calls it, and — the part that matters — when the preferred engine returns nothing or fails, it moves to the next one and records which engine answered.

# backend/smartgate/modules/search/algorithm.py — source lines 119–167 (Search.search)
    async def search(self):
        import time

        self.start_time = time.time()
        backend = self.settings.resolved_backend()
        self.effective_backend = backend
        try:
            if backend == "firecrawl":
                rows = await self._search_firecrawl()
                self.result_container.extend("firecrawl", rows or [])
            elif backend == "searxng":
                rows: list[dict] = []
                try:
                    rows = await self._search_searxng()
                except Exception as e:
                    logger.warning("SearXNG failed: %s", e)
                    if self.settings.searxng_fallback:
                        self.result_container.add_unresponsive_engine(
                            "searxng", str(e)
                        )
                    else:
                        raise
                if rows:
                    self.result_container.extend("searxng", rows)
                    self.effective_backend = "searxng"
                elif self.settings.searxng_fallback:
                    logger.info(
                        "SearXNG returned no rows; falling back to DuckDuckGo Lite"
                    )
                    try:
                        fb_rows = await self._search_duckduckgo_lite()
                        self.result_container.extend("duckduckgo", fb_rows or [])
                        self.effective_backend = "duckduckgo"
                    except Exception as fb_e:
                        logger.warning("DuckDuckGo Lite fallback failed: %s", fb_e)
                        self.result_container.add_unresponsive_engine(
                            "duckduckgo", str(fb_e)
                        )
                else:
                    logger.warning(
                        "SearXNG returned no rows (searxng_fallback=false)"
                    )
            else:
                rows = await self._search_duckduckgo_lite()
                self.result_container.extend("duckduckgo", rows or [])
        except Exception as e:
            logger.warning("Search backend '%s' failed: %s", backend, e)
            self.result_container.add_unresponsive_engine(backend, str(e))
        return self.result_container

Three details matter for a research loop. The backend is resolved at run time, so one entry point serves a self-hosted metasearch engine, a managed scraping search, or a last-resort HTML endpoint without any caller knowing which. A failed engine does not raise out of the method: it is caught, recorded on the result container, and the run continues, because a round that returns partial evidence beats one that returns a stack trace. And every success path records the effective backend, so a passage can be traced to the engine that produced it — the field an audit reads later.

For the wider map of what a research stack needs — retrieval, reading, memory and the tools around them — the cluster's centre page on AI research tools owns the selection view. This page stays on the mechanism.

deep research as a category: a plan, then rounds of retrieval

"Deep research" is a family name that several vendors now use for a class of system rather than a single product, and the shared mechanism is easier to see than the branding. Every member of the family does the same three things in the same order: it commits to a plan before it retrieves, it conditions each round on what the previous round returned, and it ends on a stop condition rather than on the model's sense of completion.

The middle property is the one that costs money, so it is worth being precise. One round is not one search request; it is a small fan-out, and part of that fan-out happens inside a single backend call. The excerpt below shows the paging arithmetic of one such backend: the container works out how many result pages it needs for the number of results it was asked for, then walks them in a loop, stopping early when a page comes back short.

# backend/smartgate/modules/search/algorithm.py — source lines 236–244 (Search._search_searxng)
        per_page = 20
        want = self.search_query.max_results
        pages_to_fetch = max(1, (want + per_page - 1) // per_page)
        start_page = self.search_query.pageno

        merged: list[dict] = []
        async with httpx.AsyncClient(timeout=30.0) as client:
            for offset in range(pages_to_fetch):
                page = start_page + offset
# backend/smartgate/modules/search/algorithm.py — source lines 256–278 (Search._search_searxng)
                resp = await client.get(
                    final_url,
                    params=params,
                    headers={"User-Agent": UA, "Accept": "application/json"},
                )
                resp.raise_for_status()
                data = resp.json()
                chunk = data.get("results") if isinstance(data, dict) else None
                if not isinstance(chunk, list):
                    break
                for a in chunk:
                    merged.append(
                        {
                            "title": (a.get("title") or "").strip(),
                            "url": (a.get("url") or "").strip(),
                            "content": (a.get("content") or "").strip(),
                            "engine": "searxng",
                        }
                    )
                if len(chunk) < 5:
                    break

        return merged[:want]

Read the arithmetic rather than the HTTP. per_page is the engine's own page size, want is what the caller asked for, and pages_to_fetch is the ceiling on how many requests one round may make. That ceiling is a budget in disguise — the difference between a round that costs one request and a round that costs eight — and it is derived from max_results, not from the model's appetite. A loop allowed to decide its own depth without such a ceiling is exactly the shape that produces the runaway research bills people describe, and the fix is a configuration value rather than a smarter prompt.

The rounds compound: each round's queries are written from the previous round's passages, so a round can only be as good as the working set it reads. That is why a deep research run's two real failure modes are spending too many rounds and stopping before the weakest claim was checked — not a model that was not clever enough. Who runs several of these a day is the subject of the AI researcher role in this cluster.

what is deep research: a system, not a longer prompt

The honest one-line answer is that deep research is not a bigger context window with better manners. Three designs get confused with each other, and they differ in where the evidence lives.

Long-context QA loads a corpus you already hold into one window and asks for an answer. There is one model call, no retrieval, and any citations come from the same text the model is summarising — you cannot follow them anywhere new.

Classic RAG adds one retrieval step in front of one generation step: the query is fixed before the model sees anything, and the answer is produced once. It reads a corpus you ingested, not the live web, so freshness is a property of your ingestion job.

Deep research keeps a working set between rounds, retrieves more than once, and separates drafting from verification. Its citations point at pages it read during the run, which is what makes them checkable.

The consequence is a different failure mode. A long-context answer fails by missing something that was in the window; a deep research run fails by trusting a page it should not have. That is why the working set has one shape regardless of where a passage came from: the excerpt below is the normaliser that turns a hit from any backend into the same small record.

# backend/smartgate/modules/search/algorithm.py — source lines 206–221 (Search._normalize_hit)
    def _normalize_hit(self, item: dict, engine: str) -> dict:
        snippet = (
            item.get("markdown")
            or item.get("description")
            or item.get("snippet")
            or item.get("content")
            or ""
        )
        return {
            "title": (item.get("title") or "").strip(),
            "url": (item.get("url") or "").strip(),
            "content": snippet.strip()
            if isinstance(snippet, str)
            else str(snippet),
            "engine": engine,
        }

The record is deliberately small — title, url, content, engine — and the engine field is the one worth noticing. It means a passage can always be traced back to the backend that supplied it, which is what makes the verification stage possible at all: support for a claim becomes a record rather than a memory. Keep that record shape stable and the rest of the loop can be replaced one stage at a time without rewriting the others.

What a synthesis stage should do with those passages, and why the drafting discipline differs from an answer engine, is the subject of AI research papers in this cluster.

how to use deep research without paying for rounds you do not need

Use it deliberately, and the first decision is not which product to open but whether the question deserves a loop at all.

Scope the question to a decision. A run pays off when the answer has to be defended — a technical choice, a market position, a question with more than one camp. If the question has a single factual answer, a search call and a sentence are cheaper and no less correct.

Set the ceiling before the run. Decide the maximum number of rounds, the maximum pages read and the maximum spend, then treat a run that needs more as failed. A loop with a ceiling degrades into a smaller report; a loop without one degrades into an invoice.

Make the plan visible. A system that shows its sub-questions lets you reject a bad decomposition before it spends the budget. If the plan is hidden, the first thing you can review is the bill.

Treat citations as claims, not proof. A citation says a page was read, not that the page supports the sentence. Spot-check the two or three claims the conclusion rests on, and treat a conclusion sentence with no citation beneath it as unsupported.

Reuse the artefacts. The durable output is usually the working set — the pages and extracts — not the prose. A run that discards it makes you pay to read the same sources twice.

deep research api: what one retrieval call hands back

The API question is not whether a research endpoint exists but what one call in the middle of the loop returns. The useful unit is a retrieval call that hands back page text, because the read stage cannot do its job on a list of links alone.

The excerpt below is a retrieval backend that does exactly that: it posts a query and asks for the results already converted to markdown, with the main content only.

# backend/smartgate/modules/search/algorithm.py — source lines 180–204 (Search._search_firecrawl)
        body: dict[str, Any] = {
            "query": self.search_query.query,
            "limit": self.search_query.max_results,
            "sources": ["web"],
            "scrapeOptions": {
                "formats": ["markdown"],
                "onlyMainContent": True,
            },
        }
        async with httpx.AsyncClient(timeout=60.0) as client:
            resp = await client.post(url, json=body, headers=headers)
            resp.raise_for_status()
            payload = resp.json()

        data = payload.get("data")
        out: list[dict] = []
        if isinstance(data, list):
            for item in data:
                out.append(self._normalize_hit(item, "firecrawl"))
        elif isinstance(data, dict):
            for item in data.get("web", []) or []:
                out.append(self._normalize_hit(item, "firecrawl"))
            for item in data.get("news", []) or []:
                out.append(self._normalize_hit(item, "firecrawl"))
        return out[: self.search_query.max_results]

Three fields carry the contract. "sources": ["web"] declares which corpora are in scope, so a round's evidence is typed rather than an uncontrolled mix. "scrapeOptions" with "formats": ["markdown"] tells the backend to fetch and convert, collapsing the fetch and read stages into one call — the reason the read stage can stay cheap. "onlyMainContent": True drops navigation and boilerplate before passages reach the model: a cleaner working set now, and a smaller bill for every later round that re-reads it.

The discipline that follows is short: name the call that produced each passage, keep the response normalised, and attribute cost per call rather than per report. An API that returns only URLs pushes the fetch stage into your own code, where it is harder to meter; one that returns markdown with the source attached keeps the read stage inside a call you can see and cap.

The API surface itself — a hosted research endpoint against a gateway you call from your own loop — is worked through on the deep research API layer.

deep research agent: a loop that owns its own stop condition

The agent-shaped variant is the one where the system chooses its own next step instead of following a fixed sequence. That freedom buys adaptability on questions whose shape is not known in advance, and it costs three things that have to be built on purpose.

A stop condition that is not the model's opinion. A maximum round count, a token budget or a wall clock: something outside the model decides when the run ends. Every runaway research story is a missing ceiling.

Durable state between rounds. If the working set lives in a process, a crash loses the run; if it lives in a store, a second worker can resume it. This is the same requirement any orchestrated pipeline has, which is why research agents and pipeline tooling keep converging on the same shape.

Degradation instead of failure. An agent that stops because one engine failed wastes everything it already read. The excerpt below is the last backend in the fallback chain — an HTML endpoint parsed directly — and its job is to keep a round alive when the preferred engines are unavailable.

# backend/smartgate/modules/search/algorithm.py — source lines 282–297 (Search._search_duckduckgo_lite)
        import httpx
        from lxml import html

        kl = _duckduckgo_lite_region(self.search_query.lang)
        form: dict[str, Any] = {"q": self.search_query.query}
        if kl:
            form["kl"] = kl

        async with httpx.AsyncClient(timeout=15.0) as client:
            resp = await client.post(
                "https://lite.duckduckgo.com/lite/",
                data=form,
                headers={"User-Agent": UA},
            )
            resp.raise_for_status()
            doc = html.fromstring(resp.text)

The interesting part is not the parsing but the position: this backend exists so that a round has a floor. An agent whose retrieval is a single API has a single point of failure; one whose retrieval is a chain can return a smaller, older or noisier set of passages and still let the verification stage judge whether they are good enough. Degradation has to be designed in, because by the time you need it the run is already underway.

The agent-shaped variant, and the orchestration it needs wrapped around the loop, is the subject of deep research agents in this cluster.

gemini deep research and the other hosted systems

Hosted systems are the visible end of this category. The names change — Gemini deep research, ChatGPT's research mode, Perplexity's research runs — but the mechanism underneath is the one this page has been taking apart: a planner, a retriever or browser that reads live pages, a working set, a synthesis pass and a citation list.

Two things are worth taking from the hosted pattern even if you never call one. First, these products made the loop legible: the user sees the plan, the sources and the report, which turned a rumour into a reviewable artefact. Second, they priced it by availability rather than by work — a limited number of runs per month — which hides the real variable, the number of rounds and pages a question actually needs.

The infrastructure question is what remains when the same loop moves inside your own product: there is no subscription to bound the cost, only a retrieval call per round, a read per page and a verification pass per draft, each of which has to be metered and capped. The mechanism does not change when the packaging does; the billing surface does.

This page does not treat any of these names as a target phrase for us. They are vendor product names, and the honest use of them is comparison, not competition for the same query.

openai deep research and the hosted pattern

The hosted pattern is worth naming on its own, because it is the shape most teams meet first: a single product surface that plans, browses and writes, with a run counter attached to a plan. Its appeal is that the whole loop is someone else's to operate. Its limit is the same as any hosted abstraction — you cannot see or tune the middle.

Three costs hide inside that surface, and naming them is what makes a build-versus-buy decision honest. Round cost: the number of retrieval and read calls a question needs is not fixed by the question, so a plan that limits runs rather than work is cheap on easy questions and expensive on hard ones. Evidence access: the report and its citation list are usually what comes back, not the passages, so reusing the reading for a second question means running the loop again. Model coupling: the planner, the reader and the writer are chosen together, so improving one stage of your own stack does not improve the run.

Where the answer has to be embedded in a product rather than read by a person, the same loop is usually assembled from primitives you operate, so that each stage is metered and can be replaced on its own — the difference between calling a research product and running a research capability.

How SmartGate fits

SmartGate is the layer this page's loop calls when the research runs inside your own product. The primitives map onto the stages one for one: smart_search issues the retrieval call, smart_fetch reads a page and returns it as markdown, smart_dedup and smart_context_gate keep the working set small between rounds, smart_memory carries the facts a later run may assume, smart_budget_guard is the stop condition that is not the model's opinion, and smart_pipe wires them into the research, read and remember templates. They are sold as five capabilities driven by seven algorithm primitives rather than as seven separate products.

The plan table is the ceiling a loop is allowed to spend against: monthly token caps of 2M, 20M, 100M and 200M+, requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and 2, 10, 30 or effectively unlimited keys per team. Reconcile those with the number of rounds your questions actually need before you commit to anything — the pricing page is the authoritative table. To see the loop on real traffic, start free on the 2M-token tier and read the per-call record on your own questions.

The part of this cluster this page does not cover is the review that has to be reproducible as a whole — how sources are searched, screened and cited as one unit — which is the job of AI literature review.

Frequently Asked Questions

Is deep research AI the same as a RAG pipeline?

No, though they share the retrieval step. A RAG pipeline fixes the query before the model runs and retrieves once; a deep research loop plans, retrieves in several rounds, keeps a working set and checks its citations. The difference is who chooses the next query, and whether the answer is verified rather than merely generated.

Why is a deep research run so much slower than a chat answer?

Because it is many calls, not one. Planning, several retrieval rounds, a page read and extract per source, a synthesis pass and a verification pass each add latency in sequence, and the slowest stage is usually fetching and reading pages. The latency is the price of the evidence trail.

Can deep research run over a private corpus instead of the web?

Yes, and the loop is unchanged; what changes is the retriever. A private index replaces the live web search call, while the working set, the rounds and the citation check stay the same. Freshness then comes from your ingestion schedule rather than from the web.

How do I keep the cost of a run predictable?

Set the ceiling outside the model: a maximum number of rounds, a maximum pages-read count and a spend cap, then measure against a per-call record. A loop without a ceiling is the only version whose cost is genuinely unbounded, because the number of rounds is not fixed by the question.

Does a longer context window remove the need for retrieval rounds?

No. A larger window lets you carry more text into one call, but it does not decide which pages to read, keep evidence between rounds, or check a citation against a source. Those are stages, not window size, and a bigger window performs none of them.

Limitations

This page is a mechanism map, not a benchmark. No hosted system is ranked, because the right shape depends on the question and on where your evidence already lives; the loop's stages are named and their trade-offs stated, and nothing here claims that any particular product implements a stage well.

The code shown is five windows into a single file — our own retrieval container — so it illustrates the retrieval and read stages and says nothing about the planning, synthesis or verification stages of any system. Those come from public sources, and every excerpt is a window rather than a whole file: the branches around each cut are described in prose instead of shown.

Degradation is a design choice with a real cost. A fallback that returns older, noisier or partial evidence can lower answer quality, so a run should record which engine answered and treat a fallback result as lower confidence rather than as an equivalent one.

The plan figures quoted here are operational limits, not a feature list, and they change with the plan; a spending decision should read the current pricing table rather than this page. Nothing here is legal or research-integrity advice: a citation a system produces is a claim about a page, and the reader still owns the judgement about whether that page is trustworthy.

Sources

Method note

The code above was cut directly out of the slice body returned by SmartGate's slice API and re-asserted byte-for-byte as a substring of that body before publication; the first line inside every fence records the file and the exact source lines. The slice matcher pinned all eight sections of this page, and every one of them resolved to the same asset — a retrieval container — which is why the page cuts five distinct regions of that one file, each labelled with its own source range, instead of repeating the same class eight times. The provenance table below lists each excerpt with the measured phrase that pinned it. No batch identifiers, auction data or internal hosts appear here, so nothing on the page has to be checked back against a build log to be trusted.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)
1 deep research ai Search backend/smartgate/modules/search/algorithm.py 119–167 rule A L2 → slot-proof 5d88cd6aed7c
2 deep research Search backend/smartgate/modules/search/algorithm.py 236–244, 256–278 rule A L2 → slot-proof 5d88cd6aed7c
3 what is deep research Search backend/smartgate/modules/search/algorithm.py 206–221 rule A L2 → slot-proof 5d88cd6aed7c
4 deep research api Search backend/smartgate/modules/search/algorithm.py 180–204 rule A L2 → slot-proof 5d88cd6aed7c
5 deep research agent Search backend/smartgate/modules/search/algorithm.py 282–297 rule A L2 → slot-proof 5d88cd6aed7c

Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before publication. 8 of 8 sections pinned, 0 abstentions, 0 misses.