SmartGateSmartGate

AI Deep Research: Commission, Check, and Keep a Run

AI deep research is a multi-round run: you hand a system one question, it plans, searches, reads and returns a cited report over several minutes. The part that decides whether that report is useful is not the model — it is the brief you write before the run, the acceptance check you apply to the report, and the record you keep afterwards. This page is that operating discipline, in that order.

Short answer: AI deep research is a multi-round run: you hand a system one question, it plans, searches, reads and returns a cited report over several minutes. The part that decides whether that report is useful is not the model — it is the brief you write before the run, the acceptance check you apply to the report, and the record you keep afterwards. This page is that operating discipline, in that order.

Key takeaways

  • An ai deep research system is a run, not a page: it has a start, a plan, a set of rounds and a stop reason, and every one of those is something you can specify before you spend anything.
  • The brief is the highest-leverage artefact. A run given a scope, a freshness window, a source policy and a budget returns something checkable; a run given one bare sentence returns something fluent.
  • Acceptance is a separate job from commissioning: the person who wrote the brief should not be the only person who reads the report.
  • Keep the record — the brief, the sources touched, the claims and their spans, the stop reason. It is the only part of a run that compounds into the next one.
  • Write the brief for one real question and run the acceptance check against a vendor's own sample report before you pay for a subscription.

ai deep research: the run is the unit of work

An ai deep research system answers a question it was not handed the sources for. There is no index to read, so it decides what to look for, goes and looks, and then decides whether what came back is enough — which is why a run takes minutes rather than milliseconds, and why the output is a report with citations rather than a paragraph. That shape is what this page treats as a unit of work; how the tools that offer it compare is a separate decision, mapped on the research tool landscape.

What matters operationally is that the run is a unit of work. It is commissioned, it executes, it is accepted, and it leaves a record. Those four moments each have an owner and an artefact, and treating them as one act — type a question, trust the answer — is where the disappointments come from. The model's quality is fixed when you buy it; the brief, the acceptance check and the record are yours, and they are where a team pulls a month of value out of the same subscription a colleague got nothing from.

Three artefacts carry the lifecycle. The brief is written before the run and says what would make the answer usable. The acceptance check is applied after it and says whether this particular report clears the bar. The record is kept when the run is done and says what was asked, what was read and what was concluded. Everything below is those three; where the category itself sits next to a search engine or a chat answer is the cluster's category definition.

grok deep research: one run shape under every brand

Grok deep research, ChatGPT deep research and Perplexity deep research are not three different machines; they are three front doors onto the same run shape. A planner drafts sub-questions, a retrieval layer fetches and normalises sources, a writer drafts the report, and a stop condition — steps, wall clock, or a token ceiling — ends the loop. What differs between the named systems is not the shape but the terms of access: where the run is started, whether it can be scheduled, what the per-run limits are, what happens to the material you send, and what the report gives you back in structured form.

That is why this page does not rank them. The differences that decide a purchase are account-level questions — can this run be triggered by code, is my data used for training, how long is the report kept — and the answers change faster than any review. The authority for any named system is that vendor's current documentation, not a comparison written this month.

What is portable is the lifecycle. The brief you would write for a grok deep research run is the same brief you would write for any other: the question, the scope, the freshness window, the source policy, the budget and the acceptance owner. Move to a different vendor and the brief survives almost unchanged, because it describes the job rather than the tool. The report format changes, the run record changes shape, and the acceptance check keeps working.

deep research tool: the brief you write before the run

The brief is the difference between a research run and an expensive guess. It is short — half a page — and every field in it removes a class of unusable output.

Brief field What it states What it prevents
Question one sentence, with the decision the answer feeds a report that answers a neighbour of your question
Scope what is in and what is out, named explicitly coverage of the next topic over
Freshness window the oldest source that may count a confident answer built on a three-year-old snapshot
Source policy the classes allowed and the classes banned an answer you cannot cite in your own document
Budget rounds, wall clock, or a spend ceiling a run that keeps reasoning until someone notices
Output contract the shape you will read it in — sections, a table, a length a wall of prose you must re-edit before use
Acceptance owner the person who will judge the report the run being accepted by the person who commissioned it

Two fields do more work than they look like they do. The freshness window is the only field that forces the system to prefer a recent source over a canonical one, and it is the field most often left implicit — which is how a run returns a well-cited answer to a question that has moved on. The acceptance owner is a control, not a formality: when the person who wrote the brief is also the only person who reads the report, the run is being graded by its own author, and the failure is invisible. The stages that sit between the brief and the report — planning, retrieval, drafting, stopping — are the cluster's stages inside a run; the brief is what you control above them.

That is also why the brief is worth writing before the tool is chosen rather than after it: the question, the scope and the output contract describe the job, so a brief that only makes sense for one vendor is a sign it has quietly turned into a feature list. Written once, in the language of the decision it feeds, the same brief fits the second run and the fifth tool, and the acceptance check you build on top of it does not have to be rewritten when the vendor changes.

ai deep research tool: accepting the report against a check

A returned report is a draft until someone accepts it, and acceptance is a check against the brief rather than a feeling about the prose. Five questions clear it or send it back.

  1. Does it answer the question in the brief? Read the question, then the conclusion. A run can produce an excellent report on an adjacent question and still be useless.
  2. Do the citations resolve, and do they say what the sentence says? Open a sample — the load-bearing claims, not the easy ones — and read the span, not just the title. A link that resolves is not the same as a link that supports the claim.
  3. Is the scope respected? Anything inside the banned classes, or outside the freshness window, is a defect even if it reads well.
  4. Is the uncertainty stated where it exists? A report with no hedges anywhere is more suspect than one that names what it could not establish.
  5. Does it say what it did not look at? The sources the run could not reach — sign-ins, paywalls, material published since the run — are the honest limits, and a report that omits them reads as complete when it is not.

The acceptance check is also where you catch the failure mode that survives good prompting: a report that is internally consistent but built on one view of the question, because every round was conditioned on the last. The symptom is a clean set of citations that all come from the same kind of source. When the underlying unit of work is a single paper rather than a question, the same discipline runs per source, and that is the job of working a single paper.

A useful habit is to accept in writing: a one-line verdict — accepted, accepted with caveats, or sent back — attached to the record turns a private judgement into something the next reader can see, and shows when it is the brief that needs the edit rather than the model.

deep research ai tool: the run record you keep

The record is the artefact almost no one keeps, and the only one that compounds. It is cheap to write because the run already produced every part of it; it is valuable because it turns a one-off answer into something a later run, a reviewer or an auditor can start from.

Record field Why it is kept
The brief, verbatim so the output can be judged against what was actually asked
Run parameters: engine, model, date, limits so a surprising report can be reproduced or explained
Sources touched, with the spans used so a claim can be checked without re-running the search
The claims and their acceptance status so the graded ledger, not the prose, is the durable output
The stop reason so "it stopped at the limit" is never mistaken for "it finished"
The cost and the wall clock so the next brief can be sized against real numbers

The field that earns its place most often is the stop reason. A run that ends because it hit the step cap has produced a partial answer, and a partial answer honestly labelled is a usable input; the same answer presented as complete is a liability. Recording the stop reason is also what makes a run budget meaningful: without it, "the run cost more this time" has no explanation, and the next brief is sized by guesswork.

Keep the record where the people who will reuse it can reach it, in a form a script can read — a row per run, with the report as an attachment rather than the record. A record that lives only as a bookmarked chat thread is not a record; it is a screenshot.

The record is also what makes a run reviewable by someone who was not there. A colleague who reads only the report has to take its citations on faith; a colleague who reads the record can see what was searched, what was ruled out and where the run stopped — which is the difference between inheriting an answer and inheriting the reason for it, and the reason the record is worth more than the report on the second use.

perplexity deep research api: when the run is a call

The lifecycle does not change when the run moves behind an API; the artefacts change shape. A perplexity deep research api call, or any hosted equivalent, is the same plan-and-retrieve loop started by code instead of a person, and it is worth choosing the trigger deliberately because it changes what the brief can be.

When a run is a call, the brief becomes parameters — the question, the allowed sources, the depth, the time ceiling — and the acceptance check becomes assertions. "The citations resolve" turns into a script that opens each returned URL and checks it is live; "the scope is respected" turns into a filter over the returned domains. This is the improvement that pays for the move: a check a person runs by hand on one report runs automatically on a thousand, and the acceptance owner becomes a rule rather than a schedule.

What does not transfer is the judgement. An automated acceptance check can confirm that a citation resolves and that no banned domain appears; it cannot decide whether the report answers the question, which stays with a person on the runs that matter. The division worth drawing in code is therefore between the mechanical checks, which run on every call, and the human read, which runs on the subset your budget says you can afford. The endpoint shapes, pagination and error contracts that sit underneath this are the cluster's programmatic research endpoint; what stays on this page is that the brief and the record have to exist in code before the call is worth automating.

chatgpt deep research api: the hosted-run defects that reach the record

Hosted runs fail in a handful of characteristic ways, and the useful thing about all of them is that they leave a trace in the record if the record has the right fields. Naming them is how acceptance turns into something you can check rather than something you hope.

Silent truncation. A run collects more sources than the report lists, and the report presents a bounded set as the set. The record catches it when it keeps the sources touched, not just the sources cited, because the gap between the two is the truncation. Freshness drift. The run reaches a conclusion from an older snapshot because its retrieval favoured the canonical text; the freshness window in the brief is the only thing that pushes against it, and the record's run date is what makes the drift visible after the fact. A budget stop dressed as a finish. The run hit a per-run ceiling and returned what it had, phrased as though it were complete; the stop reason in the record is the field that refuses to let the two be confused. Undisclosed reach. The run could not see a class of source — signed-in, licensed, or published after it ran — and the report does not say so, which the acceptance check's last question is there to catch.

None of these is a reason to avoid a hosted run; they are the reason the record exists. A chatgpt deep research api run with a brief, a stop reason and a kept source list is auditable. The same run with only the report is a paragraph nobody can check, however good it reads.

deep research agent github: keeping the lifecycle in your own repo

When you own the loop — a deep research agent github project you run yourself, or a hosted run you drive through your own service — the three artefacts become files, and the discipline becomes much easier to enforce because a repository remembers what a chat window forgets.

The mapping is direct. The brief becomes a small checked-in document, one per recurring question, versioned so a change to scope is a change you can see. The run record becomes an append-only file, one row or one small file per run, holding the parameters, the sources, the claims and the stop reason. The acceptance check becomes a script in the repository that reads the record and fails loudly on a banned domain, a dead citation, or a missing stop reason. The payoff is that the parts of the lifecycle that are mechanical — freshness, scope, citation liveness, budget — stop depending on a person remembering to look.

What stays out of the repository is the part that is genuinely the vendor's: the loop itself. Building your own deep research agent is a different project with a different cost profile, and the cluster's write-up of the agent architecture behind a run is where that decision belongs. For most teams the honest split is a hosted run and a repository that holds the brief, the record and the check — the loop is rented, the discipline is owned.

The commissioning checklist, in the order it saves money

Written as a sequence, because each step makes the next one cheaper to get right and none of the first three costs anything.

Step What you write The question it settles
1 The brief for one real question what would make this answer usable, before any run
2 The acceptance check, as five questions who reads the report, and against what
3 The record template, with the stop-reason field what survives the run, and in what form
4 One small paid run against the brief whether the tool clears your bar at all
5 The scheduled runs, once the check is written which questions are worth asking every week

The order is a dependency rather than a preference: the acceptance check reads the record, the record is only useful if the brief defined what good looks like, and no purchase decision can be made honestly before one run has been graded. A team that buys the subscription first and writes the brief last is paying for fluent reports it cannot use.

Where SmartGate fits

SmartGate is the intelligence layer a research practice runs on, and the reason to put a gateway between your runs and the open web is the same as the reason for the record: the brief, the limits and the evidence all need one place to live. The five capabilities map onto the lifecycle directly — Research is smart_search plus smart_fetch, which is the retrieval a run's rounds stand on; Context is smart_dedup plus smart_context_gate, which is what keeps a long run's own context from crowding out contradicting evidence; Memory is smart_memory, the team store that lets a later run start from what an earlier one established; Control is smart_budget_guard, which checks, counts and records against a hard ceiling; and Pipeline is smart_pipe, which orders research, read and remember into one callable job so the record is written without a second integration.

That matters most for the field this page keeps returning to: the stop reason. A gateway that sees every tool call can bound a run in code rather than in a prompt, and the audit row it writes — caller, route, token count, outcome — is the same object the budget reads and the record keeps. The plan table is about operational limits rather than features: monthly token caps of 2M, 20M, 100M and 200M+, requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and a team holding 2, 10, 30 or unlimited keys. Billing shares only in what the platform saves, and the pricing page is the authoritative table — read those numbers there rather than here.

How duplicate sources are collapsed, read from our own dedup module

Deep research fails in a predictable way: the same claim arrives five times from five URLs and looks like five sources. Here is how ours is actually collapsed, read on 2026-10-07 from backend/smartgate/modules/dedup/.

  • Exact duplicates are handled first, and cheaply. A record layer groups rows by an exact key while preserving first-occurrence order, so identical items never reach the expensive step. Ordering is preserved on purpose: a "keep the first" rule is only reproducible if the first is well defined.
  • Similarity is a vector problem, and the index says so. Duplicates that are not identical are found with an index that holds its vectors in memory and delegates querying to a separate backend — the split that makes a small deployment practical without a vector database service.
  • The encoder is an interface, not a model. The module defines an encoder protocol, so what turns text into vectors is replaceable. That is what lets the same deduplication logic serve text today and something else later — and it is also the place where a silent quality change can enter.
  • The shipped research chain draws the line at 0.85 similarity. That is a dial with two failure modes, and both are worse than they look: turn it down and distinct sources collapse into one (your report loses its corroboration), turn it up and the same page under two URL forms passes twice (your report gains false corroboration).
  • The honest limit: deduplication reduces redundancy, it does not establish truth. Two independent pages agreeing is evidence; two copies of one page are not, and only the tool knows which it kept.

How to get started

The first three steps need no purchase at all.

  1. Write the brief for the one question you would ask this week if a research run were free. Question, scope, freshness window, source policy, budget, output contract, acceptance owner.
  2. Write the five-question acceptance check, and name the acceptance owner. This is the step that makes the rest testable.
  3. Run one vendor's own sample report through the check, honestly. If it fails the check, no subscription fixes that.
  4. Run one real question, keep the record, and grade the report against the brief.
  5. When the discipline holds, start free with an MCP-speaking client and point the retrieval step at the gateway, then read the pricing page once your real call pattern tells you which tier you need.

Frequently Asked Questions

Is ai deep research just a chatbot with search?

No. A chatbot with search makes one retrieval and one answer; a deep research run plans, retrieves over several rounds, and returns a report with citations. The difference that matters operationally is that the run has a plan and a stop reason you can inspect, which is what makes the brief and the record worth writing.

Do I need a deep research tool, or is a search engine enough?

Most questions are closed by one good retrieval and a careful read, in which case a run adds minutes and budget for nothing. The run earns its cost when the answer needs several dependent steps that cannot be known in advance, and when you can state what would make the answer usable before it starts. If you cannot write the brief, you are not ready for the run.

How long should a run be allowed to take?

Short enough that hitting the ceiling is a design choice rather than an incident. Set a step cap, a wall clock and a spend ceiling together, and record which one stopped the run. A partial answer with its stop reason stated is a usable input; the same answer presented as complete is the failure this page is written against.

Which of the named systems should we use?

Whichever one lets you trigger and record the run on your terms — a callable run, a structured output and a retention policy you can live with beat a marginally better report. That is an account question rather than a review question, and the vendor's current documentation is the only authority for it. The brief you write is portable between them.

What should the record actually contain?

The brief verbatim, the run parameters, the sources touched with their spans, the claims and their acceptance status, the stop reason, and the cost and wall clock. Keep it where a script can read it, one row per run. The report is an attachment to the record, not the record.

How many rounds should a run take?

As many as the question needs and no more — "needs" is bounded by the acceptance check, not by the model's appetite. Most useful runs finish in a handful of rounds, and a run still retrieving without adding a distinct source has run out of novelty; recording which limit stopped it keeps "long" from being mistaken for "thorough".

Limitations

This page describes an operating discipline, not a product ranking: no named system is scored, because the right choice turns on account-level terms that change faster than any review, and the discipline works the same whether a run is hosted, self-built or started by hand.

The lifecycle does not fix a thin source base. If the answer depends on material behind a sign-in, on a licensed database, or on a document published minutes ago, no brief and no record will retrieve it, and the honest outcome is a clearly-scoped partial answer with its stop reason stated. The acceptance check is a filter, not a guarantee: it tells you when a report fails a stated test, and it cannot certify that a passing claim is true, only that the cited span supports it as written.

The plan figures quoted above are operational limits that move with the plan, and a compliance decision should read the pricing page rather than this page. This page carries no code excerpt, and the reason is recorded in the Method note below: no line-numbered or implementation-level claim is made about any system, including ours.

The five-question check is only as strong as the person who wrote it: a check that asks only whether citations resolve will pass a report that cites the wrong span fluently. The remedy is a second reader on the runs that carry weight.

Sources

  • OpenAI's deep research announcement — Introducing deep research, for the multi-round research shape and the plan-and-budget concerns behind it.

  • Google's research-agent rollout — Gemini Deep Research, as the consumer-facing form of the same run shape.

  • Anthropic's engineering note — Building effective agents, for the workflow-versus-agent split and why an explicit stop condition is the line between the two.

  • The ReAct paper — arXiv:2210.03629, for the interleaved reason-and-act loop a research run formalises, and Chain-of-Verification — arXiv:2309.11495, for checking a drafted claim against its span.

  • The Model Context Protocol specification — modelcontextprotocol.io/specification, as the wire contract between a run and the tool layer it calls.

  • The Wikipedia article on deep research and the Hacker News and Reddit threads that rank for the head phrase were read as evidence of what the searcher already knows, not as sources for the discipline above.

  • Demand figures on this page are this project's own paid measurements, recorded in its search_volume.json and research_brief.md; product behaviour was read read-only from the product source at the revision this project's slice run recorded.

  • The section "How duplicate sources are collapsed, read from our own dedup module" is our own implementation, read on 2026-10-07 from backend/smartgate/modules/dedup/records.py, modules/dedup/index.py and modules/dedup/utils.py (origin/main). It states only what those files state.

Method note

This page carries no code excerpt, and the reason is a measurement rather than an omission. The slice matcher pinned 8 of 8 sections for this project (0 abstentions, 0 no-slice misses), but every pin resolved to the same general-purpose asset: a web-search container class whose name collides with ordinary search vocabulary across the product codebase. One generic symbol standing behind eight different section topics is not section-level evidence, so the page is written from sources instead — a pinned generic would have given it the shape of a verified article with none of the substance.

Product claims above (the primitive set and the plan limits) were read read-only from the product source at the revision this project's slice run recorded, and the plan figures were re-checked against the live pricing page on 2026-10-03. The demand figures come from this project's own paid keyword measurement, not from a third party. No code, batch fingerprints, auction data or internal hosts appear on the page, so there is nothing here that has to be asserted verbatim.