What Is Deep Research? Definition, Scope, and Trust
Deep research is an AI workflow that turns one question into a multi-step investigation: it plans sub-questions, searches, reads many live sources over several minutes, and returns a written report with citations. It is not a search engine, not a single chat answer, and not an autonomous agent. Judge any report by whether its citations open and its sources hold up.
Short answer: Deep research is an AI workflow that turns one question into a multi-step investigation: it plans sub-questions, searches, reads many live sources over several minutes, and returns a written report with citations. It is not a search engine, not a single chat answer, and not an autonomous agent. Judge any report by whether its citations open and its sources hold up.
Key takeaways
- The definition is about the process, not the product. Deep research is the loop of plan, search, read, deduplicate, compress and cite. A vendor name tells you which implementation you met, not what the workflow is.
- Depth is measured in steps and sources, not in minutes. A run that read forty sources and merged them is deep research; a run that read none is a long answer wearing the same label.
- The three boundaries are the whole definition. Search returns ranked links, a one-shot model answer synthesises from memory, an agent takes actions. Deep research reads live sources on your behalf and stops at a report.
- Citations are the product. A sentence you cannot trace to a URL you can open is not a finding; it is a claim, and the report's value is exactly the difference between the two.
- A run has a predictable timeline — restate, plan, search, fetch, dedupe, compress, synthesise, cite — and knowing that order is how you debug a bad report instead of rewriting the prompt.
- Next step: before you trust or ship any deep-research output, run the four checks in the trust section below against one report you already have open.
What deep research is: the definition, and the line it draws
Deep research is a way of answering a question by spending time and reading instead of by predicting. The output is not a link list and not a paragraph; it is a short document whose sentences are meant to be traceable back to the pages the run actually opened. Everything else in this article follows from that one property.
Four things have to be true at once for the label to be honest, and each of them is checkable after the fact rather than taken on faith.
It plans before it searches. A single query is a lookup. A deep-research run first turns your question into a set of sub-questions — the dimensions a good answer would have to cover — and then searches for those. If a report cannot tell you what it looked for, it did not plan.
It reads live sources. The answering model may know a great deal already, but a deep-research run is supposed to fetch pages at run time, so that a fact published last week is eligible to appear. This is the difference between a memory of the web and a reading of it.
It reconciles what it read. Searching ten times returns the same page ten times, and returns two pages that disagree. Deduplication and compression are not housekeeping; they are what lets a thousand fetched paragraphs become a report that fits in a reading session without the reader seeing the same sentence three times.
It cites. Every non-obvious claim carries a source, and the source is a place you can go. A footnote that does not resolve is worse than no footnote, because it launders a guess as a finding.
None of that requires a particular vendor. The workflow is a shape, and the shape is old — an analyst does the same thing over three days with a browser and a notebook. What the AI version changes is the cost and the wall-clock time of the middle steps, not the definition. Where you buy those middle steps is a different question, and it belongs to the category map on AI research tools.
What deep research is not: a search engine, a chat answer, or an agent
Most confusion about the term comes from three neighbours that share its vocabulary. The boundaries are worth stating precisely, because "deep research" is now used for all three by people who mean very different things.
It is not a search engine. A search engine ranks documents against a query and returns links. The work of reading them is yours. It answers "where might the answer be", not "what is the answer". Deep research inverts that: the reading is the product, and the links are evidence attached to prose.
It is not a one-shot chat answer. A single model turn answers from what it already holds, compressed into its parameters at training time. That is often good enough and it is fast, but it has two structural properties you should not confuse with research. First, it cannot cite a page it did not fetch, so any citation is reconstructed rather than observed. Second, it has no way to know what changed after its training cut-off, so a confident answer about a fast-moving field is a memory, not a reading. A deep-research run costs minutes precisely because it refuses both shortcuts.
It is not an autonomous agent. This is the boundary people get wrong most often, because the two share a tool-calling loop. An agent acts: it writes files, calls APIs, sends messages, changes state in the world, and can keep going until a stop condition fires. A deep-research run reads: search, fetch, deduplicate, compress, write a report, stop. It has side effects on nothing but the report. The moment a research loop is allowed to send, deploy or purchase, it is an agent and it needs the controls that word implies — permissions, budgets, an audit trail — which is the subject its own page covers on deep research agent.
The practical test is one sentence: if the worst thing a runaway run can do is waste money and produce a bad document, it is research. If it can change the outside world, it is an agent, and the report is the least of your concerns.
How to use deep research: one run, start to finish
Here is the timeline a run actually follows, and it is the same order whether the runner is a hosted product or a script you wrote. Reading it in this sequence is what turns a mysterious result into a diagnosable one.
1. Restate the question. The runner rewrites your prompt into an explicit, bounded question — usually with a scope, an audience and a time frame. This step is where vague prompts go wrong: ask "tell me about the market" and the restatement will pick a market for you, silently. Read the restatement first; if it disagrees with what you meant, stop and fix the question before spending the run.
2. Plan the sub-questions. The question is decomposed into the few dimensions an answer needs. This is the step that separates research from search, and the plan is usually visible in the output. A plan with three sub-questions produces a report with three sections of uneven depth; six produces a more even one.
3. Search, across more than one engine. Each sub-question becomes several queries, and the queries go to search backends — often more than one, because a single engine's index is a single point of failure. This is the first step that touches the code, and it is where a run's coverage is decided. The excerpt below is the search entry point of a real implementation: it resolves which backend to use, starts the clock, and routes one engine's hits into a shared result container.
4. Fetch and read. The ranked hits are not the answer; the pages behind them are. The runner downloads each candidate and converts it to text, which is where most of the wall-clock time goes — a run takes minutes because it is waiting on the open web, not because the model is thinking.
5. Deduplicate and compress. Forty pages contain the same paragraph forty times. The pile is reduced to distinct claims and then compressed to fit whatever context the synthesis step has. A run that skips this step does not fail loudly; it produces a long report that repeats itself, which is the most common quality complaint about the genre.
6. Synthesise. Only now does the model write. The interesting property is that it writes from retrieved text rather than parametric memory, which is why the citations can be real.
7. Cite and return. Claims are attached to their sources and the report is handed back.
The step that people most often misunderstand is the third, because it is invisible in the output. A search step has to sit in front of several engines, each with its own response shape, and its whole job is to make that difference disappear before the next step sees anything. Two windows out of that step show the pattern.
# backend/smartgate/modules/search/algorithm.py — source lines 119–128 (python)
async def search(self):
import time
self.start_time = time.time()
backend = self.settings.resolved_backend()
self.effective_backend = backend
try:
if backend == "firecrawl":
rows = await self._search_firecrawl()
self.result_container.extend("firecrawl", rows or [])
# backend/smartgate/modules/search/algorithm.py — source lines 206–221 (python)
def _normalize_hit(self, item: dict, engine: str) -> dict:
snippet = (
item.get("markdown")
or item.get("description")
or item.get("snippet")
or item.get("content")
or ""
)
return {
"title": (item.get("title") or "").strip(),
"url": (item.get("url") or "").strip(),
"content": snippet.strip()
if isinstance(snippet, str)
else str(snippet),
"engine": engine,
}
Two things are happening there. The first window is the entry point: it resolves a backend from runtime
settings, records which backend actually answered, starts the clock, and extends one shared result
container rather than returning an engine-specific structure. The second window is the more important
half: whatever the engine called its text — markdown, description, snippet or content — the hit
that leaves this step has exactly four fields, and one of them is the URL. That normalisation is the
mechanism behind the trust section further down: a downstream step never learns which engine answered,
but every claim it carries still has an address attached to it.
deep research ai: what the model adds, and where it breaks
It is worth being precise about the model's contribution, because "AI" is doing a lot of unearned work in the marketing of this category. The model is not what makes the research deep. It is what makes the research affordable.
Its first contribution is decomposition: turning a question into sub-questions is a language task, and it is the step a fixed script cannot do well. Its second is compression: deciding which of forty overlapping paragraphs carry the claim is again a language judgement, and doing it mechanically loses the point. Its third is synthesis: writing from the retrieved pile rather than from memory.
Each of those is also a failure mode, and they fail in ways that look like success.
The plausible citation. A model asked to attribute claims will produce an attribution whether or not the source supports it, and a well-formed citation is more dangerous than a missing one because it survives a casual skim. Any practitioner in this field will tell you the same thing: verifying the citations is not optional.
The confident synthesis. Two sources disagree and the report silently picks one, or averages them, without saying that the field is split. Fluency is not the same as accuracy, and a report that reads well is not more likely to be right.
The invisible coverage gap. A run cannot cite what it did not find, so it cannot tell you that it missed the most important document in the area. An answer built from six mediocre pages will be as confident as one built from the six best, and the difference is invisible from the outside. What the model brings to this category, and its limits in more detail, is the frame of deep research AI.
The honest summary is that the model makes the middle of the workflow cheap and leaves the ends — asking the right question, and checking the answer — exactly as expensive as they were.
deep research tools: the four layers of the category
When people say "deep research tools" they are usually pointing at four different layers, and comparing across them is why tool discussions go nowhere. Knowing which layer a product occupies tells you what you are actually buying.
Layer 1 — the hosted product. A finished application: you type a question, wait, and read a report. The model, the search, the reading and the writing are somebody else's problem. This is the layer most people meet first.
Layer 2 — the open framework. A repository you run yourself that implements the loop and lets you swap the model and the search backend. You gain control of the pipeline; you inherit its operations.
Layer 3 — the research primitives. Search, fetch, dedupe, context compression and memory, exposed individually so a builder can assemble a loop inside their own product. This is the layer where the question stops being "which tool is best" and becomes "which pieces do I not want to rebuild".
Layer 4 — the corpus layer. Where the sources come from when the open web is the wrong input: paper databases, internal document stores, licensed archives. A literature-grounded run is only as good as the index behind it, and that layer has its own page for the task it serves, AI literature review.
The layers stack. A product at layer 1 is usually built out of primitives at layer 3 and a corpus at layer 4, and a framework at layer 2 is a way of wiring the same pieces. Comparing a product to a primitive is how a category gets a reputation for being overpriced — you are pricing a finished outcome against a component.
deep research alternatives: what you are really choosing between
The word "alternative" implies a like-for-like substitution, and the useful question is the one it hides: alternative for what job, at what volume, and with what tolerance for a wrong answer.
The human analyst. Still the reference point. An analyst brings judgement about what matters, asks better follow-up questions and knows the domain's landmarks. What deep research beats is throughput and cost per report; it does not beat judgement, and it will not tell you which of the six documents is the load-bearing one. AI researchers is where that division of labour is worked out role by role.
A search engine plus an afternoon. For a question whose answer lives on one page, a search engine and twenty minutes of reading is strictly better: faster, cheaper and correct. Deep research pays off when the answer is spread across many pages, when the pages disagree, or when you would have to read a hundred to find the ten that matter.
A retrieval pipeline over a fixed corpus. If your sources are a known, bounded set — your own documents, a licensed archive — a retrieval-augmented system over that corpus will beat an open-web run on both precision and cost. Reading the open web is a feature when the answer is out there, and a liability when it is not.
A survey or a literature review. When the question is "what does the literature say", the right instrument is a review process, not a web run: coverage has to be systematic and the inclusion criteria have to be stated. That is a different task with a different contract.
Choosing between them is not a preference about AI. It is a question about where the answer lives and whether a fluent wrong answer is expensive. Where it is, the substitute is not another tool but a process.
chatgpt deep research and perplexity deep research: the two starting points
Because two product names dominate the searches around this definition, it is worth saying plainly what they are and are not. Both are implementations of the workflow above at layer 1: you ask a question, they plan and search and read, and they return a cited report a few minutes later. Neither is the definition, and the definition survives either of them being replaced.
What such products share is the shape — a plan, a set of fetched sources, a synthesis, a citation list. They differ in the details that matter operationally: which search index they read, how the plan is exposed, how the citations are presented, how long a run may take, and what plan you need to start one. Those are purchasing and workflow questions rather than definitional ones, and they change faster than the definition does.
The practical advice is the same for both. Read the plan before you read the report. Open three citations at random and check them against the sentences that cite them. And re-run one question twice: if the two reports agree on the facts and differ only in emphasis, you are looking at a stable answer; if they disagree on facts, you are looking at a search-coverage problem and neither report is finished.
The name you start with is a starting point, not a category. The category is the workflow, and merging that observation with the paper-side workflow is what turns it into a habit — the connection AI research papers covers for the readers who work from the literature.
How to judge a deep research report: clickable citations and verifiable sources
The whole reason to spend minutes instead of seconds is that the output is supposed to be checkable. Here is the checklist, in the order that costs you the least time.
Open the citations. Not all of them — three, chosen at random from different sections. A citation that does not resolve, or resolves to a page that does not contain the claim, is the single strongest signal that the report was assembled rather than researched. Do this before you read the prose; it takes ninety seconds and it decides whether the rest is worth your time.
Check that claims map to sources, not that sources exist. A report can carry twenty real URLs and still attach the wrong one to a sentence. The test is to read one sentence and ask whether the page behind its footnote actually says that. The answer is usually yes and occasionally no, and the no is what you were paying for.
Look for recency where it matters. A field that moves monthly needs sources from this month. If a report about a fast-moving topic cites nothing from the last two quarters, the run read an archive, not a web.
Check whether it says "the sources disagree". A report that never admits uncertainty is either about a settled question or has quietly resolved a conflict it should have surfaced. The better reports say which part of the answer is contested, and that sentence is worth more than the confident ones around it.
Ask what it did not cover. Compare the sub-questions in the plan against the sections in the report. A dimension that was planned and then dropped is where the answer is thinnest, and the runner will not flag it for you.
That checklist turns a report from an oracle into an instrument, which is all it ever was. It is also short enough that there is no excuse for skipping it — the failure mode of this genre is not that people trust bad reports, it is that they trust good-looking ones without opening three links.
The shapes compared in one table
| Shape | What it returns | Reads live sources | Cites | Can act on the world | Fits when |
|---|---|---|---|---|---|
| Search engine | ranked links | no | n/a | no | the answer is on one page |
| One-shot chat answer | prose from memory | no | not reliably | no | the question is settled and general |
| Retrieval over a fixed corpus | passages from your documents | no (your corpus is static) | yes, internally | no | the sources are known and bounded |
| Deep research | a cited report | yes | yes | no | the answer is spread across many pages |
| Research agent | the same, plus actions | yes | yes | yes | a follow-up step must change state |
The row that decides most projects is the fourth. If the answer is spread out, disagreeing, or expensive to be wrong about, the workflow is the right shape and the vendor is a detail. If a follow-up step has to write, send or buy something, you are on the last row and you need the controls that word implies, not a better prompt.
Frequently Asked Questions
Is deep research the same as a long AI answer?
No. Length is not the distinction. A long answer is still one pass over what the model already holds. A deep-research run reads sources at run time, which is why it takes minutes and why its citations can be opened. If nothing was fetched, nothing was researched.
Do I need to be technical to use it?
No. The hosted products make it a text box and a wait. Being technical buys you control over which index is read and where the report goes next, which matters when the output has to live inside a product rather than in a chat window.
How long should a run take?
Long enough to fetch pages: a few minutes for a normal question. A run that returns in seconds did not read anything, and a run that takes half an hour has usually wandered. Duration is a diagnostic, not a quality score.
Why does the same question give different answers twice?
Because the web is not the same twice and the plan is not fixed. Two runs sample the search results at different moments, and a busy query returns a different top set each time. Stable facts usually survive; unstable ones are exactly the ones worth checking by hand.
Can I use the output as a source in my own writing?
Not directly. Treat it as an index: open the citations, read the primary sources, and cite those. The report tells you where to look, and quoting it instead of its sources is how a real finding turns into a hearsay claim.
Limitations
- A run cannot find what it did not search for. Coverage is bounded by the queries the plan chose and the index behind them, and the report has no way to tell you about the document it missed.
- Confident prose is not evidence. The workflow improves the provenance of a claim, not its truth. A citation that resolves is necessary and not sufficient; you still have to read the page.
- Paywalled and login-walled sources are usually skipped. The open web is what a run can read, so literature behind a subscription is under-represented, which matters most in the fields where the best sources are paywalled.
- Long runs cost tokens, and cost scales with reading, not with value. A run that fetched fifty pages and used three paid for all fifty. The budget belongs on the run, not on the invoice.
- This page is a definition, not a buyer's guide. It deliberately does not rank products or quote feature lists, because those move faster than the definition and belong to the evaluation pages in this cluster.
- The demand figure is a snapshot. One country, one window, one endpoint; the head phrase is a small, low-competition query, and a number measured today is not a forecast.
Sources
- OpenAI's announcement of the workflow and its stated method — Introducing deep research.
- Google's description of the same shape as a product — Gemini Deep Research.
- A widely used open implementation of the loop, useful for reading how a run is wired — dzhng/deep-research.
- The encyclopaedic summary of how the term entered general use — ChatGPT Deep Research.
- Anthropic's engineering note on when a fixed workflow beats an open agent — Building effective agents.
- For the substrate that makes an affordable run possible, the plan table this page links to: pricing.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|---|---|---|---|---|---|
| 1 | how to use deep research | Search |
backend/smartgate/modules/search/algorithm.py |
119–128, 206–221 | rule A L2 → slot-proof | 5d88cd6aed7c |
Method note
The matcher pinned 8 of 8 sections for this page (0 abstention(s),
0 no-slice verdict(s)). All eight sections resolved to the same symbol — one Search
container, matched because every keyword carries a search or research token — so that one symbol is
quoted once, in the run-timeline section, rather than repeated eight times. The other sections are
written from sources, as the house rule requires when a section has no distinct slice of its own, and
no code is transcribed anywhere on the page. Every product claim was read from the product source at
the revision pinned in this project's pipeline results, read-only; the plan figures were re-verified
against the live pricing page on 2026-10-02. The quoted block is cut from the slice body and re-asserted
against it byte for byte before publication, and no batch fingerprints, auction data or internal hosts
appear in the text.