The AI Researcher Workflow: What to Automate, What to Keep
An ai researcher's project moves through seven steps — frame, search, screen, close-read, note, write, review — and only the middle three are safe to hand to a model. Search, screening and note-taking can be automated because a person still opens the raw material; close reading, synthesis and the final review stay human, because a fluent wrong answer there becomes a retracted paper.
Short answer: An ai researcher's project moves through seven steps — frame, search, screen, close-read, note, write, review — and only the middle three are safe to hand to a model. Search, screening and note-taking can be automated because a person still opens the raw material; close reading, synthesis and the final review stay human, because a fluent wrong answer there becomes a retracted paper.
Key takeaways
- A model is a fast filter, not an author: use it to rank, group and summarise, and never to assert a finding you have not opened the source for.
- The delegable steps share one property — their output is a pointer to material a person can still check.
- The dangerous steps share the opposite property — their output is an assertion, and a fabricated citation reads exactly like a real one.
- Give every automated step a source link and a human sign-off row, or the time saved in screening is spent repairing references later.
- Start by automating the head of the funnel, and keep a person at the tail of it.
The ai researcher's project as a seven-step pipeline
Every project — a PhD chapter, a competitive scan, a two-week industry question — moves through the same seven steps: frame the question, search, screen the results, close-read the survivors, take notes, write, and review. The steps differ in cost but not in order, and the order is what makes the automation question answerable. A model can be dropped into any of them; the useful question is which step's output a person can still check without redoing the work.
The dividing line is not "creative versus mechanical", because every step here is mechanical in part. It is whether the step's output is a pointer or an assertion. Search and screening return pointers: a paper, a passage, a candidate list. A person opens the pointer and decides. Close reading and synthesis return assertions: "this result holds", "these three papers agree". Once a model has stated an assertion in its own words, the error and the truth read the same, and the checking stops.
This page maps that line onto a real working day. It does not rank tools — the selection framework for that is on AI research tools — and it does not cover the mechanics of an automated research loop, which sit with the sibling pages linked below. It answers one question: for each of the seven steps, hand it over, keep it, or split it, and what breaks when the call is wrong.
Step 1 — framing the question stays human
Framing is the first step where a model's speed is a liability. A usable research question is narrow enough to answer and broad enough to matter, and it is defined as much by what you deliberately exclude as by what you ask. A model asked to "refine this question" will obligingly broaden it, add sub-questions, and return a scope that sounds thorough and cannot be finished in the time you have.
The defensible use is adversarial rather than generative. Write the question yourself, then ask the model to attack it: what evidence would count against this, what am I assuming, and which of these sub-questions is really a different project. That use is safe because the output is a list of challenges a person evaluates; the model is not being asked to state what is true, only to find the weak joint.
One habit makes this step cheap later. Write the inclusion and exclusion criteria before the search, not after. Screening rules invented mid-review are how a review drifts, quietly, toward whatever the search engine happened to surface first — and the drift is invisible in the final draft.
Step 2 — search: the layer a machine can own
Search is the first step that is safe to automate end to end, because a search returns pointers and a reader can still ignore them. This is also the step teams over-buy: retrieval is not the hard part, deciding what to keep is. What an automated search step must do is boring and testable — send a query to more than one source, normalise the results into one shape, and keep going when a source fails.
The failure behaviour is the part worth reading, because a retrieval layer that returns an empty list when its first backend is down is worse than no automation: it looks exactly like "nothing exists on this topic", and that false negative is what sends a literature review back to the drawing board three weeks later.
# backend/smartgate/modules/search/algorithm.py — source lines 119–167 (three-backend fan-out with recorded fallback)
async def search(self):
import time
self.start_time = time.time()
backend = self.settings.resolved_backend()
self.effective_backend = backend
try:
if backend == "firecrawl":
rows = await self._search_firecrawl()
self.result_container.extend("firecrawl", rows or [])
elif backend == "searxng":
rows: list[dict] = []
try:
rows = await self._search_searxng()
except Exception as e:
logger.warning("SearXNG failed: %s", e)
if self.settings.searxng_fallback:
self.result_container.add_unresponsive_engine(
"searxng", str(e)
)
else:
raise
if rows:
self.result_container.extend("searxng", rows)
self.effective_backend = "searxng"
elif self.settings.searxng_fallback:
logger.info(
"SearXNG returned no rows; falling back to DuckDuckGo Lite"
)
try:
fb_rows = await self._search_duckduckgo_lite()
self.result_container.extend("duckduckgo", fb_rows or [])
self.effective_backend = "duckduckgo"
except Exception as fb_e:
logger.warning("DuckDuckGo Lite fallback failed: %s", fb_e)
self.result_container.add_unresponsive_engine(
"duckduckgo", str(fb_e)
)
else:
logger.warning(
"SearXNG returned no rows (searxng_fallback=false)"
)
else:
rows = await self._search_duckduckgo_lite()
self.result_container.extend("duckduckgo", rows or [])
except Exception as e:
logger.warning("Search backend '%s' failed: %s", backend, e)
self.result_container.add_unresponsive_engine(backend, str(e))
return self.result_container
The excerpt above is a single search container that fans out to three backends and records which one actually answered. Two decisions are worth copying into any research stack. First, failure is not an exception that escapes the function — the container logs the failure, records the engine as unresponsive (add_unresponsive_engine), and continues, so a partial result set still reaches the researcher with a note about what was missing. Second, the fallback chain is short and explicit (a configured engine, then a second, then a third), which means a person can read the whole control flow in one screen and know which source answered a given query.
Where that becomes a research workflow rather than plumbing, the reasoning layer on top of retrieval is deep research, and the orchestrated, multi-step version of it is deep-research agents. Both sit downstream of the same primitive: reliable, repeatable search.
Step 3 — screening: the biggest time win, and its limit
Screening — deciding which of four hundred hits deserve a read — is where a model saves the most wall-clock time, and where a quiet error compounds fastest. At its core, title-and-abstract screening is a classification task: does this result meet the inclusion criteria. Framed that way, a model is genuinely useful, because the cost of a wrong "maybe" is one wasted look and the cost of a wrong "no" is a paper you never see again.
Two rules keep the automation honest. Screen in two passes — a cheap recall-oriented pass that keeps anything plausibly relevant, then a precision pass on the shortlist — and measure both passes against a hand-labelled sample of about fifty items, so the model's false-negative rate is a number rather than a hope. And keep the model's reason next to each decision, because "why was this excluded" is the question a co-author, or a reviewer, will ask about your method.
The limit is that screening decisions are not facts to be averaged. A model that excludes a paper with an articulate justification will convince you; a model that includes one is only spending your time. That asymmetry is why the machine's reject list deserves more auditing than its keep list, and why the reject list is the one that should be seeded from the model but signed off by a person.
The review-level version of this step — query design, deduplication across databases, and the screening record a reader can audit — is worked through in full on AI literature review.
Step 4 — close reading: where the model must not decide
Close reading is the step where a model may summarise but must not conclude. Asking a model to "read this paper" is fine; asking it to "tell me what this paper found" is where fabricated and real findings become indistinguishable. The model will return a confident paragraph with the right vocabulary, and that paragraph may describe a paper that does not say any of it.
The workflow that survives makes the model a retrieval aid inside a text you already trust. Point it at the methods and results sections and ask for the specific sentence that supports a claim, then read that sentence. Quote-then-read is slower than summary-then-trust, and it is the only version where an error is visible before it reaches a draft. When the material is a single paper, the paper-level techniques — figure and table extraction, claim-to-passage linking — are collected on reading an AI research paper.
There is a second, subtler failure. A model summarising ten papers will silently merge them, attributing study A's effect size to study B because both were "about the same thing". Nothing in the prose flags the merge; the sentence is grammatical, specific, and wrong. The defence is structural: one source per claim in your notes, with a synthesis sentence that names the set it summarises rather than gesturing at "the literature".
Step 5 — note-taking and the citation trap
Note-taking looks administrative and is actually where the citations are won or lost. The safe pattern is a note that stores a quotation plus a stable identifier — DOI, arXiv id, page and paragraph — and never a paraphrase the model composed from memory. Paraphrase is fine when it sits directly beneath the quotation it came from, because then the check is a glance instead of a re-read of the source.
The citation trap is specific and well documented. A model asked to produce references will produce plausible ones: correct author-looking names, correct journal-looking titles, correct DOIs that resolve to a different paper or to nothing at all. This is not a rare glitch. In Mata v. Avianca an attorney filed a brief citing cases that did not exist, and the court sanctioned the filing — the first widely reported case of fabricated citations reaching a court record, and the reason a citation is now treated as a claim that needs its own verification step.
The defence costs seconds and belongs in the tooling, not the willpower: every reference in your draft links to a page you have opened, and a script checks that each DOI resolves before the draft leaves your hands. Treat "the model gave me a citation" the same way you would treat a citation from an anonymous forum post — as a lead, not a fact.
Step 6 — writing: what to delegate, what to own
In a research draft, the model is a good line editor and a poor author. The prose that carries your contribution — the argument, the ordering of evidence, the sentence that says what the result means — is the part a reader is assessing you for, and delegating it produces the flat, confident register that reviewers now recognise on sight.
What is safe to delegate is the scaffolding: turning bullet points into readable paragraphs, tightening a methods section, generating a first-pass abstract that you then rewrite, and checking that your terminology is consistent. What must not be delegated is the mapping from evidence to claim, because that mapping is the research. A useful discipline is to write every claim sentence yourself and let the model only ever improve the sentences around it.
A second discipline is to keep a "claims I cannot yet support" list. When the model drafts a transition that asserts something you have not verified, the sentence is neither accepted nor deleted; it is parked on that list until a source turns up or the claim is dropped. Making the gap explicit is what stops a smooth draft from accumulating unexamined assertions.
Step 7 — review: the last human gate
Review is the step that makes the other six honest. Before the draft leaves, one person who did not write it opens a sample of sources and checks three things: the claim is present in the source, the citation points at the right paper, and the number matches. A review that only reads the prose catches style and misses everything that matters, because the whole failure mode of AI-assisted research is prose that is fluent and wrong.
Sampling is enough if it is random and recorded. Checking every claim is the ideal and rarely survives a deadline; checking a random ten percent of claims, plus every claim that carries a number, plus every citation the model supplied, catches the defects that scale. Record which claims were checked and who checked them, so the review is an artefact and not a memory.
The final gate is a lane rule as much as a reading rule. A model may draft, rank and summarise; it may not be the last reader of a factual claim. Put a person between the model's output and the publication, and the workflow changes from "fast and fragile" to "fast and checkable".
Where the workflow backfires
The failures above are not random. They cluster at the step boundaries, and each has a signature worth naming before it costs you a draft.
| Failure | Where it enters | Why it is invisible | The cheap check |
|---|---|---|---|
| Fabricated citation | Note-taking and writing | A fake reference looks like a real one | Resolve every DOI before the draft leaves |
| Merged findings | Close reading across papers | One grammatical sentence, two sources | One citation per claim, no exceptions |
| Silent broadening | Framing the question | The scope still sounds specific | Attack the question before searching |
| Missed paper | Screening | The exclusion reads as decisive | Audit the reject list, not the keep list |
| Plausible summary | Close reading | The vocabulary is right, the finding is not | Quote the supporting sentence, read it |
The pattern is that every failure is a correct-looking output at a step where a person was supposed to check a pointer and instead trusted an assertion. That is why the delegation question is not "how good is the model" — the same model is safe at search and dangerous at synthesis — but "does this step still leave a pointer for a human to open".
A handoff map for one project
Written as a sequence, because each step's automation decision is cheaper once the previous one is fixed.
| Step | Hand to a model? | What the person still owns | Cost of getting it wrong |
|---|---|---|---|
| 1. Frame | No — only to attack the question | The question, the scope, the exclusions | An unanswerable or drifting project |
| 2. Search | Yes, end to end | The query, and reading the failure report | A false "nothing exists" |
| 3. Screen | Yes, with an audited reject list | The inclusion criteria, the sign-off | A missed paper you never see |
| 4. Close-read | Summarise, never conclude | The supporting sentence and its reading | A plausible, wrong finding |
| 5. Note | Format, not paraphrase from memory | Quotation plus identifier per note | A fabricated citation in the draft |
| 6. Write | Scaffolding and line edits only | The evidence-to-claim mapping | A flat draft that is yours in name only |
| 7. Review | No | The final factual read and sign-off | A fluent error that ships |
Read top to bottom, the column that says "hand to a model" is not a recommendation to automate everything it touches — it is the boundary beyond which the next column stops working. Automate step 2 without the failure report and step 3 screens a truncated list; automate step 5 without the quotation and step 7 inherits citations nobody can check.
How this fits SmartGate
The workflow above needs one thing from infrastructure: a retrieval and processing layer whose failures are visible, so the person at steps 5 and 7 can trust what reaches them. That is the layer SmartGate provides. Five capabilities — Research, Context, Memory, Control and Pipeline — sit over the primitives an agent calls: multi-engine search and URL-to-Markdown fetch for the search step, de-duplication and context handling for the reading step, memory for the notes, and budget control and pipeline orchestration so a long research loop cannot run away.
Applied to this page's steps: the search primitive feeds step 2 and reports which sources answered, so a partial result set is a fact rather than an empty list; the de-duplication and context primitives serve step 4, where ten papers repeat each other and the cost is paying a model to read the same paragraph ten times; memory carries step 5 across projects; and budget and pipeline orchestration are what make an automated loop safe to leave alone. The point is not to replace the researcher at the steps this page keeps human — it is to make the machine's output checkable at exactly those steps.
The plan table sets the operational limits rather than the features: monthly token caps of 2M, 20M, 100M and 200M+, MCP requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and 2, 10, 30 or unlimited keys per team. Pro starts at $18 per month. The pricing page is the authoritative table, and the login page is where a researcher can start with the free tier to test the search and de-duplication primitives on a real reading list before committing to a plan.
If you build agents rather than use them directly, the programmatic entry point for the same loop — the one that lets your own pipeline own steps 2 and 3 — is the deep research API.
Frequently Asked Questions
Can an AI researcher write a literature review on its own?
No, and the failure is specific rather than general. A model can run the search and the first screening pass well, but the review's contribution is the synthesis, and a synthesis the model states in its own words cannot be checked sentence by sentence without redoing the reading. Use it to build the candidate list and keep a person on the synthesis.
Where do AI-assisted researchers fabricate citations most often?
At the moment a reference is written from memory, which is usually during note-taking and drafting. The model is not retrieving the paper when it prints a citation; it is generating a plausible string. The fix is mechanical: every citation is added from a page you have open, and every DOI is resolved before the draft leaves.
Is it safe to let a model exclude papers during screening?
Only if the reject list is audited. A model's exclusions read as decisive even when they are wrong, and a wrongly excluded paper is invisible in your final draft. Keep the model's recall high, check its reject decisions against a labelled sample, and let a person sign off on what gets dropped.
How many sources should a model summarise at once?
Few enough that each claim keeps one citation. Summarising ten papers in one pass is where findings get merged and attributed to the wrong study. Summarise one paper per claim, or ask the model to name the source for every sentence, and treat any sentence that cites "several studies" as unfinished work.
What is the single rule that keeps this workflow honest?
A model may rank, group and summarise, but a human opens the source before a claim enters the draft. Every step that respects that rule is safe to automate; every step that quietly violates it is where a fluent, confident, wrong paragraph enters the paper.
Limitations
This page describes a workflow, not a validated protocol, and it makes no claim about how much time any particular team will save. The seven-step split is a model of how research projects move, not a standard; a review conducted to PRISMA or Cochrane requirements will have its own mandatory steps and its own audit trail, and those requirements override anything here.
The safety boundary — pointers are delegable, assertions are not — is a heuristic, not a proof. It fails in both directions: a model can be wrong about a pointer, and a careful human-in-the-loop synthesis can still be wrong. What the boundary buys is visibility: it keeps every automated decision attached to material a person can reopen, so the worst case is corrected rather than discovered after publication.
Finally, the plan figures quoted above are operational limits read from the pricing page, not a feature comparison, and they change with the plan. A decision that depends on retention or rate limits should read the current table rather than this page.
Sources
- Anthropic's engineering note on workflow and agent patterns — Building effective agents, for the workflow-versus-agent split and where a human decision belongs.
- The PRISMA 2020 statement — prisma-statement.org/prisma-2020-checklist, the reporting standard a screening record is audited against.
- The Cochrane Handbook for Systematic Reviews of Interventions — training.cochrane.org/handbook, for screening and data-extraction discipline.
- OpenAI's announcement of its research agent — Introducing deep research, the shape of an automated multi-source research loop.
- Google's Gemini Deep Research announcement — next-generation Gemini Deep Research, a second implementation of the same loop.
- The Model Context Protocol specification — modelcontextprotocol.io/specification, for how a retrieval tool is described to an agent.
- Mata v. Avianca, Inc. — Wikipedia, the case that made fabricated citations a sanctionable risk rather than a curiosity.
- Demand figures in this page are our own measurements: DataForSEO Google Ads, United States, 12-month window, recorded in this project's
search_volume.json. Product behaviour and the plan table were read from the product source at the revision pinned in this project'spipeline_results.json, read-only.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|---|---|---|---|---|---|
| 1 | ai researcher | Search |
backend/smartgate/modules/search/algorithm.py |
119–167 | rule A L2 → slot-proof | 5d88cd6aed7c |
Method note
The page quotes one code excerpt, and the choice is deliberate: the slice matcher pinned 6 of 8 sections for this page (0 abstention(s), 2 no-slice verdict(s)), but all pinned sections resolved to the same symbol — the search container. Rather than repeat one implementation six times to match the section count, the page shows that implementation once, at the step it belongs to, and writes the remaining sections from sources. The excerpt was cut from the slice body and re-asserted byte-for-byte, so the code is the file's, not a transcription. No batch fingerprints, auction data or internal hosts appear on this page; the section keyword quoted above each heading comes from this project's own paid measurement run.
Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before publication. 6 of 8 sections pinned, 0 abstentions, 2 misses.