AI Research Agent: Where Automation Should Stop
An ai research agent should automate retrieval, extraction, de-duplication and drafting, and should stop short of anything a reader will treat as a fact until a human has verified it. The practical line is four gates — citation verification, sampled fact checks, budget approval, and the audit trail — plus one rule: retrieval may degrade silently, but a conclusion may not be shipped from a…
Short answer: An ai research agent should automate retrieval, extraction, de-duplication and drafting, and should stop short of anything a reader will treat as a fact until a human has verified it. The practical line is four gates — citation verification, sampled fact checks, budget approval, and the audit trail — plus one rule: retrieval may degrade silently, but a conclusion may not be shipped from a degraded run without a named reviewer.
Key takeaways
- Draw the line at conclusions, not at steps. Searching, fetching, de-duplicating and drafting can run unattended; a number, a quotation or a citation that a reader will repeat should pass a named human first.
- Gate four things and leave the rest alone: citation verification, a sampled fact check, budget approval, and the audit trail. Gating everything is the same as gating nothing, because a reviewer under load rubber-stamps.
- Make each gate readable from the record, not from memory. The provenance fields a reviewer needs are the source URL, the engine that produced it, the fetch time, and the reviewer's name — all four, in one row.
- Set the stop conditions before the run starts. A maximum step count, a token ceiling and a wall clock are what turn an autonomous research agent into one you can operate and bill.
- Start with one output class: pick the sentence a reader is most likely to repeat, make it require sign-off, and log who signed. Then widen the gate only where the record shows a real miss.
For the reader who runs a research team rather than an agent framework, the argument in one paragraph: an ai research agent is unusually good at the parts of research nobody enjoys — running one query across several sources, fetching the pages, removing duplicates, and producing a first draft — and unusually bad at knowing when it is wrong. The failure that costs money is not slowness; it is a confident sentence built on a degraded source that nobody checked. So the design question is not "how autonomous can it be" but "which outputs must a human own". The AI research tools category, surveyed end to end, exists to answer the first half of that question; this page answers the second half with four gates and the record that makes each gate auditable.
ai research agent: where automation should stop
An ai research agent is a loop that decides what to search, fetches what it finds, and writes up an answer. The code below is the dispatcher at the front of that loop. It resolves a backend from settings, records which backend actually answered, and handles the two supported paths — a direct Firecrawl call, or a SearXNG call wrapped in its own error handling. The governance-relevant detail is the pair of assignments to effective_backend and the add_unresponsive_engine call in the failure branch: when a source fails, the agent does not stop, it downgrades and keeps going.
That is the right behaviour for retrieval and the wrong behaviour for a conclusion. A reader does not care which search backend answered a question; a reader cares whether the sentence in front of them is true. The automation boundary this page argues for sits exactly there. The agent may choose, fetch, degrade and retry on its own — that is the work it is good at — but the moment its output is going to be repeated as fact, the run has to hand a human something checkable: the claim, its source, and how the source was reached. Everything that follows is one of the four gates that make that hand-off real, and the record that makes each gate auditable instead of ceremonial.
# backend/smartgate/modules/search/algorithm.py — source lines 119–143 (class Search: search dispatcher)
async def search(self):
import time
self.start_time = time.time()
backend = self.settings.resolved_backend()
self.effective_backend = backend
try:
if backend == "firecrawl":
rows = await self._search_firecrawl()
self.result_container.extend("firecrawl", rows or [])
elif backend == "searxng":
rows: list[dict] = []
try:
rows = await self._search_searxng()
except Exception as e:
logger.warning("SearXNG failed: %s", e)
if self.settings.searxng_fallback:
self.result_container.add_unresponsive_engine(
"searxng", str(e)
)
else:
raise
if rows:
self.result_container.extend("searxng", rows)
self.effective_backend = "searxng"
ai research tool: the provenance a reviewer checks first
Every candidate a research agent produces has to arrive in one shape before a reviewer can compare it, and that shape is a design decision with governance consequences. The normalizer below collapses four different upstream payloads — Firecrawl markdown, a description, a snippet, a generic content field — into the same four keys, and the fourth key is the interesting one. engine records which backend produced the row, which is what separates a finding from a fallback.
For a reviewer, that field is the first thing to read. A citation is only as strong as its source URL, and a source URL is only as strong as the pipeline that produced it. If the row came from the primary backend, the reviewer verifies the claim against the page. If it came from a fallback, the reviewer is not merely checking a quotation, they are deciding whether the source is admissible at all — because a degraded path is also a path whose failure mode the run never surfaced to the caller. The same discipline applies when the input is a document rather than a web page: extracting from an AI research paper means recording the section and page a quotation came from, not just the paper, so a second reader can confirm the sentence in the same place.
# backend/smartgate/modules/search/algorithm.py — source lines 206–221 (class Search: _normalize_hit)
def _normalize_hit(self, item: dict, engine: str) -> dict:
snippet = (
item.get("markdown")
or item.get("description")
or item.get("snippet")
or item.get("content")
or ""
)
return {
"title": (item.get("title") or "").strip(),
"url": (item.get("url") or "").strip(),
"content": snippet.strip()
if isinstance(snippet, str)
else str(snippet),
"engine": engine,
}
ai web research: confirming what the agent actually fetched
An ai web research step returns a snippet, not a page, and the difference is the whole gate. The fetch below asks the backend for markdown with only the main content, normalizes each returned item, and then slices the list down to the requested maximum. Two deliberate reductions sit in those five lines: the page is reduced to its main content, and the result set is reduced to a count. Both are correct for a first pass and both are disqualifying for a fact, because neither the reader nor the reviewer ever sees the bytes the claim was built from.
The review rule is therefore blunt: a reviewer verifies a web claim against the URL, not against the snippet the agent kept. A snippet that supports a sentence is evidence that the page might support it; only the page itself is evidence that it does. When the fetch is part of a larger summarization pass — the shape a literature review takes — the reviewer should sample at least one claim per source and open the URL, because the failure this catches is not a hallucinated citation, it is a real citation whose surrounding context says the opposite.
# backend/smartgate/modules/search/algorithm.py — source lines 180–204 (class Search: _search_firecrawl)
body: dict[str, Any] = {
"query": self.search_query.query,
"limit": self.search_query.max_results,
"sources": ["web"],
"scrapeOptions": {
"formats": ["markdown"],
"onlyMainContent": True,
},
}
async with httpx.AsyncClient(timeout=60.0) as client:
resp = await client.post(url, json=body, headers=headers)
resp.raise_for_status()
payload = resp.json()
data = payload.get("data")
out: list[dict] = []
if isinstance(data, list):
for item in data:
out.append(self._normalize_hit(item, "firecrawl"))
elif isinstance(data, dict):
for item in data.get("web", []) or []:
out.append(self._normalize_hit(item, "firecrawl"))
for item in data.get("news", []) or []:
out.append(self._normalize_hit(item, "firecrawl"))
return out[: self.search_query.max_results]
ai research assistant: what the reviewer, not the model, decides
The polite word for an agent that answers research questions is "assistant", and the assistant's weakest moment is the one worth gating. When the primary search returns no usable rows, the branch below does one of two things: if fallback is enabled it quietly retries a different backend and records that the first one was unresponsive; if fallback is disabled it logs a warning and returns an empty container. In neither case does the caller learn that the answer rests on a weaker source — the container simply has fewer rows, or rows from somewhere else.
That is where a human gate earns its place. The reviewer's job is not to re-run the whole search; it is to decide whether an answer assembled from a degraded retrieval step is good enough to publish, re-run, or kill. Three questions make that decision in under a minute. Did the run fall back at all, and if so how far? Does the answer change if you drop the fallback rows? And is any single load-bearing sentence supported by more than one independent source? A yes to the last question is what lets the answer leave the gate. The wider product question — where an assistant stops being a tool and starts being a deep research system — is a different page's job; the gate here only needs the three answers on the record.
# backend/smartgate/modules/search/algorithm.py — source lines 144–160 (class Search: empty-result fallback)
elif self.settings.searxng_fallback:
logger.info(
"SearXNG returned no rows; falling back to DuckDuckGo Lite"
)
try:
fb_rows = await self._search_duckduckgo_lite()
self.result_container.extend("duckduckgo", fb_rows or [])
self.effective_backend = "duckduckgo"
except Exception as fb_e:
logger.warning("DuckDuckGo Lite fallback failed: %s", fb_e)
self.result_container.add_unresponsive_engine(
"duckduckgo", str(fb_e)
)
else:
logger.warning(
"SearXNG returned no rows (searxng_fallback=false)"
)
ai researcher: who signs off on a claim
A research agent's run state is created per instance and dies with it, which is fine for computation and fatal for accountability. The constructor below shows the whole of that state: a result container, a settings object, a start time, an effective-backend string. Nothing in it names a person, and nothing in it survives the process. A run that produced a published claim and then vanished has, from an auditor's point of view, produced nothing at all.
So the fourth gate is not a check on the agent, it is a check on the human, and it needs three fields the code above cannot supply: who approved the claim, when, and on what evidence. The practical form is a review row that points at the run's record and carries the reviewer's name — the same shape as a code review, and for the same reason. It also changes what "an ai researcher" means inside a team: the person is no longer the one who runs every query, they are the one whose name is on the claims, and the how an AI researcher works split between delegation and ownership depends on exactly that distinction. Without a name on the row, the split does not exist.
# backend/smartgate/modules/search/algorithm.py — source lines 102–117 (class Search: run state)
class Search:
"""搜索容器 — SearXNG JSON / Firecrawl API / DuckDuckGo Lite。"""
effective_backend: str = ""
def __init__(
self,
search_query: SearchQuery,
settings: SearchRuntimeSettings,
):
self.search_query = search_query
self.settings = settings
self.result_container = ResultContainer()
self.start_time = None
self.actual_timeout = None
self.effective_backend = ""
research copilot: the line between suggestion and decision
A research copilot assembles candidates; a person decides what to do with them. The request the code below builds is a good picture of where the assembler's authority ends. It chooses which engines to query, which categories to search, how many pages to walk, and what language to use — all of which shape the candidate set — and then it stops. Nothing in it commits to an answer, and that restraint is the design principle worth copying: the copilot's output is a menu, and the menu is only as useful as the decision that follows it.
Put the review gate at the decision, not at the menu. Reviewing every candidate a copilot produces is unbounded work and will be skimmed; reviewing the decision the copilot is about to drive is bounded and consequential. Concretely, the gate sits on four decision types — a claim that will be published, a recommendation that will be acted on, a spend that will be committed, and a conclusion that will be filed as prior work. For each, the reviewer needs the copilot's candidate set plus the reason the chosen item won. A copilot that cannot show why one candidate beat another has not made a decision, it has made a suggestion, and the two are not interchangeable.
# backend/smartgate/modules/search/algorithm.py — source lines 223–250 (class Search: _search_searxng request)
async def _search_searxng(self) -> list[dict]:
"""与 Firecrawl searxng_search 一致:GET .../search,format=json。"""
import httpx
raw = self.settings.searxng_url.strip()
cleaned = raw[:-1] if raw.endswith("/") else raw
final_url = f"{cleaned}/search"
engines_param = self.settings.searxng_engines
if self.search_query.engines:
engines_param = ",".join(self.search_query.engines)
categories_param = self.settings.searxng_categories
per_page = 20
want = self.search_query.max_results
pages_to_fetch = max(1, (want + per_page - 1) // per_page)
start_page = self.search_query.pageno
merged: list[dict] = []
async with httpx.AsyncClient(timeout=30.0) as client:
for offset in range(pages_to_fetch):
page = start_page + offset
params: dict[str, Any] = {
"q": self.search_query.query,
"language": self.search_query.lang,
"pageno": page,
"format": "json",
}
autonomous research agent: what may run unattended
The word "autonomous" should describe the retrieval loop, not the publishing decision. The fallback path below is a fair picture of the edge an unattended run reaches: when the preferred JSON backends are unavailable, it falls to an HTML endpoint, parses the result links, walks sibling nodes for a snippet, and returns what it found. That is a competent last resort. It is also a fragile source, reached silently, on a code path that was never exercised in the happy case — which is precisely the profile of output that should not become a published fact without a human in the loop.
What may run unattended is therefore a short list with a hard boundary. Unattended: scheduling, querying, fetching, de-duplicating, drafting, and the retry ladder between backends. Human-gated: the decision to publish, the decision to spend beyond an approved ceiling, and any claim that will be attributed to a person or a team. Two numbers make the boundary enforceable rather than aspirational — a maximum step count and a token budget — because both are set before the run and both can stop it without asking the model whether it is done. An agent that cannot be stopped by its own budget is not autonomous, it is unsupervised.
# backend/smartgate/modules/search/algorithm.py — source lines 280–297 (class Search: _search_duckduckgo_lite)
async def _search_duckduckgo_lite(self) -> list[dict]:
"""DuckDuckGo Lite HTML(POST),不依赖 Google DOM。"""
import httpx
from lxml import html
kl = _duckduckgo_lite_region(self.search_query.lang)
form: dict[str, Any] = {"q": self.search_query.query}
if kl:
form["kl"] = kl
async with httpx.AsyncClient(timeout=15.0) as client:
resp = await client.post(
"https://lite.duckduckgo.com/lite/",
data=form,
headers={"User-Agent": UA},
)
resp.raise_for_status()
doc = html.fromstring(resp.text)
deep research agent: the stop conditions a human sets
A deep research agent is defined less by how far it can go than by when it stops, and the bound below is a small, honest example. The parser reads page after page from the backend and breaks as soon as a page comes back thinner than a handful of rows: the loop stops because the source has stopped being informative, not because a model decided it was satisfied. That is the difference between a stop condition and a stopping feeling, and it is the same distinction a reviewer applies when reading the run's record afterwards.
The human sets three bounds up front and one after. Up front: a maximum number of steps, a token or currency ceiling, and a wall clock — so that the run ends with something the operator can predict. After: the review gate on whatever the run wants to publish. The architecture that makes those bounds real — the loop, the state, how steps are recorded — is the subject of the deep research agent architecture; this page only insists that the bounds exist and are readable from the record. A run whose stopping point cannot be explained after the fact cannot be reviewed, and an unreviewable research agent is a liability regardless of how good its answers are.
# backend/smartgate/modules/search/algorithm.py — source lines 262–278 (class Search: searxng result bound)
data = resp.json()
chunk = data.get("results") if isinstance(data, dict) else None
if not isinstance(chunk, list):
break
for a in chunk:
merged.append(
{
"title": (a.get("title") or "").strip(),
"url": (a.get("url") or "").strip(),
"content": (a.get("content") or "").strip(),
"engine": "searxng",
}
)
if len(chunk) < 5:
break
return merged[:want]
Where SmartGate fits
SmartGate is the place in a research pipeline where the record and the enforcement point are the same object. Every tool call through the gateway — a search, a fetch, a context compression, a memory read — is written as an audit row, and the same row is what the per-key rate limit and the per-team token budget read. That matters for the gates above because a review needs a record before it needs a policy: a reviewer who can see the source URL, the engine that produced it, and the tokens it cost can make the three decisions this page asks for in one screen, and a team that never has that row is making them from memory.
The plan limits are the operational half of the same story. Monthly token caps run 2M, 20M, 100M and 200M+ across the four tiers; requests per minute per key run 120, 300, 600 and 1200; audit-log retention is 7, 30, 90 or 180 days; and a team holds 2, 10, 30 or unlimited API keys. The retention window is the number to check first if a review or audit deadline drives you, because it decides how far back the gate can actually reach. Treat the pricing page as the authoritative table and read the current values there before committing a date.
| Layer | What it decides | Who owns the decision |
|---|---|---|
| Retrieval and fetch | Which sources are queried and pulled | Agent, unattended |
| Normalization | How a candidate is shaped for review | Agent, by design |
| Fact and citation gate | Whether a claim may be published | Named human reviewer |
| Budget and stop conditions | When a run ends | Human, set before the run |
| Audit trail | What can be re-checked later | Platform, retention-bound |
Frequently Asked Questions
Which parts of an ai research agent can run without a human?
Searching, fetching, de-duplicating, normalizing sources and drafting all run unattended. The parts that need a human are the decisions: what may be published as fact, what may be spent beyond an approved ceiling, and any claim attributed to a person or team. Automation is safe up to the point where the output stops being a draft and starts being evidence.
What counts as a fact check a reviewer has to do?
Open the source URL and confirm the claim in context, not just the snippet the agent kept. Sampling works: check every load-bearing claim and one claim per source on the rest. The check catches the case a snippet cannot show you, which is a real citation whose surrounding text argues the opposite of the sentence built on it.
How do we audit a research agent after the fact?
From a record that outlives the run and carries four fields per claim: the source URL, the engine that produced it, the time of the fetch, and the name of the reviewer who approved it. If your agent's state is per-process and dies with it, you have no audit trail yet, only logs. The audit begins when those four fields are written to durable storage you control.
Does a review gate make the agent too slow to be useful?
A well-placed gate makes it usable. The agent keeps doing the unbounded work; the human reviews a bounded set of decisions, not every candidate. Gate four things, set them to the decisions that carry consequences, and the review stays a few minutes per run. Gate everything and the reviewer will skim, which is slower than no gate because it hides the misses.
What should a deep research agent stop on?
A maximum step count, a token or currency ceiling, and a wall clock set before the run, plus the source thinning out on its own. A run that stops because a model decided it was finished has no explainable stopping point; a run bounded by numbers can be explained, billed and reviewed. Prefer bounds that exist without asking the model.
Limitations
This page is a control design, not a certification. It argues for four review gates and describes the fields each one needs; it does not claim that any particular agent, vendor or framework implements them, and it does not audit any named product's trail. The criteria are the ones a reviewer can apply in minutes, which means they are not exhaustive — they are chosen to be used.
The provenance discussion is limited by the code it quotes: the excerpts come from one search module in one product, and they illustrate the retrieval and normalization layer rather than the reasoning layer. A gate built on these fields governs what the agent fetched and how it was shaped; it does not, on its own, verify that the sentence drawn from a source is the sentence the source supports. That last step stays a human judgement, which is the point of the gate rather than a gap in it.
The plan figures above are operational limits, not a feature comparison, and they move with the plan. The retention window in particular is the number most likely to be wrong for your deadline, so it should be read from the pricing page rather than from this page. Nothing here is legal or compliance advice, and nothing here asserts that a review gate satisfies any regulatory regime.
Sources
- NIST AI Risk Management Framework — nist.gov/itl/ai-risk-management-framework, for the govern-map-measure-manage structure that a human review gate fits into.
- The Model Context Protocol specification — modelcontextprotocol.io/specification, for how a runtime tool call is described and returned, which is the boundary this page's gates sit on.
- OpenTelemetry semantic conventions for generative AI — opentelemetry.io/docs/specs/semconv/gen-ai, for the span attributes an audit row can carry instead of bespoke glue.
- Anthropic's engineering note on building effective agents — anthropic.com/engineering/building-effective-agents, for the workflow-versus-agent split that decides what may run unattended.
- Demand figures in this page are our own measurements: DataForSEO Google Ads, United States, 12-month window, recorded in this project's
search_volume.jsonandresearch_brief.md. - Product behaviour and the plan table were read from the product source and re-checked against the live pricing page at write time.
Method note
The code excerpts above were not transcribed. Each fenced block was cut out of the slice body returned by the slice API, using the source line windows recorded in this spec, and then re-asserted byte-for-byte against that body before the page was rendered; the first line inside every fence names the file and the source lines it came from, and the table below gives one row per excerpt. The section keyword quoted above each heading comes from this project's own paid measurement run, not from a third-party tool.
Product behaviour was read from the product source at the revision the slice run recorded, read-only. No batch fingerprint, auction figure, internal host or private address is written into the page, and the plan limits above are quoted from the public pricing table. The page states its pin count in digits so that the number can be checked against the matcher's own record rather than read as prose.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|---|---|---|---|---|---|
| 1 | ai research agent | Search |
backend/smartgate/modules/search/algorithm.py |
119–143 | rule A L2 → slot-proof | 5d88cd6aed7c |
| 2 | ai research tool | Search |
backend/smartgate/modules/search/algorithm.py |
206–221 | rule A L2 → slot-proof | 5d88cd6aed7c |
| 3 | ai web research | Search |
backend/smartgate/modules/search/algorithm.py |
180–204 | rule A L2 → slot-proof | 5d88cd6aed7c |
| 4 | ai research assistant | Search |
backend/smartgate/modules/search/algorithm.py |
144–160 | rule A L2 → slot-proof | 5d88cd6aed7c |
| 5 | ai researcher | Search |
backend/smartgate/modules/search/algorithm.py |
102–117 | rule A L2 → slot-proof | 5d88cd6aed7c |
| 6 | research copilot | Search |
backend/smartgate/modules/search/algorithm.py |
223–250 | rule A L2 → slot-proof | 5d88cd6aed7c |
| 7 | autonomous research agent | Search |
backend/smartgate/modules/search/algorithm.py |
280–297 | rule A L2 → slot-proof | 5d88cd6aed7c |
| 8 | deep research agent | Search |
backend/smartgate/modules/search/algorithm.py |
262–278 | rule A L2 → slot-proof | 5d88cd6aed7c |
Every fenced block above was cut from the slice body returned by the API and re-asserted against it byte-for-byte before publication. The matcher pinned 8 of 8 sections for this page (0 abstention(s), 0 no-slice verdict(s)), and every excerpt above is one of those pins.