AI Research Paper: The Five Stages and the Three Accidents
An AI research paper is a paper whose search, analysis or drafting ran with a language model in the loop. The model is genuinely strong at finding sources, compressing them and drafting prose, and it is the wrong owner of two things: the reference list and the statistics. Keep those two human-checked, and the rest of a paper gets faster without getting riskier.
Short answer: An AI research paper is a paper whose search, analysis or drafting ran with a language model in the loop. The model is genuinely strong at finding sources, compressing them and drafting prose, and it is the wrong owner of two things: the reference list and the statistics. Keep those two human-checked, and the rest of a paper gets faster without getting riskier.
Key takeaways
- A paper moves through five stages — question, method, analysis, writing, submission — and a model that is safe in one is a liability in another. The stage, not the model, decides the risk.
- Retrieval is the safest place to start: an AI research tool answers "what already exists on this question" far faster than a person, and a wrong search costs a revision, not a retraction.
- The reference list is where AI papers break. A measured audit of 178 model-suggested references found 28 that resolved to nothing at all.
- Statistics must be recomputed, never summarised. If a model wrote the sentence, a human runs the test.
- Figures and tables are read, not trusted: a vision model misreads a chart in ways that read as confidence, and a reviewer may not check the number behind the claim.
- Write the verification in before the draft, not after: one row per claim, with the source that proves it, and a named human who signed it off.
ai research paper: what the finished artefact has to contain
The phrase names a paper whose literature search, analysis or prose passed through a language model, and it is a label for the process rather than for a genre. What reaches a reader is still an ordinary paper: a question, a method, a result, a discussion and a reference list. What changed is who produced each part and how fast.
That shift is worth taking seriously because the failure modes are not the ones people expect. A model rarely writes a sentence that is grammatically wrong; it writes a fluent sentence attached to a source that does not exist, or a confident summary of a number it did not compute. The damage is not in the prose, which reads well, but in the two artefacts a reader is least likely to re-derive: the citations and the statistics. Both are checkable, and that is the whole argument of this page — the way to use a model on a paper is to make every claim it touches checkable by someone else.
It is also worth separating this page from its neighbours. This is not a literature-review page: the review is one stage with its own reporting standard, handled in the cluster. It is not a workflow page about agents that run unattended. This page follows a single paper from the question to the submission and asks, at each step, which part a model may own and which check has to survive it.
The five stages, and the one that decides whether you keep the paper
A paper is a sequence, and the sequence is stable across fields. Naming the stages makes the risk assignable, and the front of it — retrieval — is where AI research tools concentrate.
| Stage | What a model does well | The failure that reaches a reader | Who signs off |
|---|---|---|---|
| Question and scoping | map what already exists, surface adjacent work | a question that is already answered, or one nobody can test | the author |
| Method and protocol | draft the design, list the assumptions | a protocol written after the results, not before | the author |
| Analysis | write the code, run the standard test | a summary of a number nobody recomputed | the author, by re-running it |
| Writing | draft, tighten, restructure | fluent prose over a wrong or missing source | the author |
| Submission and review | format, check the checklist, rebut | a response letter that answers a different question | the author |
The stage that decides whether the paper survives is the third one, and not because analysis is hard. It is because analysis is summarisable: a model can produce a paragraph about a result it never computed, and the paragraph is indistinguishable, on the page, from one that came from a real run. A wrong search is corrected in an afternoon; a wrong statistic after review, or after a correction notice. The other stages are recoverable in proportion to how far the error sits from that third one.
ai researcher: where the model works and where the author has to
The role question has a clean answer: a model is a fast, tireless, unreliable assistant, and the reliable part of its value is the work that can be checked by inspection. The judgement that carries the author's name is the work that cannot.
Three jobs belong to the author and should not be delegated, not because a model would refuse them but because it would do them fluently. Framing the question is a judgement about which gap is worth a year, and it depends on taste and standing that no retrieval layer holds. Choosing the method is where a model will happily draft three designs and rank none of them; the ranking is the contribution. And deciding what the result means is the one part of a paper that is genuinely irreplaceable, which is why it is the part reviewers argue with.
What is left is a large, honest block of assistance: reading, locating, comparing, compressing and reformatting, plus the first draft of anything that will be rewritten anyway. The division is not technical versus non-technical; it is checkable versus not. A model that finds forty candidate papers has done checkable work — every one can be opened and rejected. A model that says the field "broadly agrees" has done uncheckable work wearing the costume of a finding. What an AI researcher is now, and how the role is taught, belongs to that cluster page; the boundary this page adds is a rule: delegate the retrievable, keep the signable.
research paper summarizer ai: reading a field without losing the thread
Summarisation is the stage a model was built for, and it is also where the first quiet error enters, because a summary is a lossy artefact and the loss is invisible. A summariser that has read two hundred abstracts and returns a tidy synthesis has done something no person can do in an hour, and has also thrown away exactly the details that decide whether the synthesis is true.
The discipline that keeps a summary usable has three parts. Keep the pointer, not just the claim: every sentence in a synthesis should be traceable to the abstract, page and figure it came from, so that the summary's value survives the first contradiction. Summarise per source, then group: asking a model for a single answer over the whole pile invites it to average away the one paper that disagrees, which is usually the interesting one. And treat the summary as a reading list, never as a result: the output is a map of where to look, and the map is checked by opening the sources it names.
This is the stage where an AI research tool earns its place, because the work is retrieval and compression over a corpus too large to read and too important to skip. The trap is treating the compressed corpus as evidence: a summary of a claim is not the claim, and a reviewer who follows it will find the original or not find it at all. The small habit worth building: after any model-written synthesis, open three of its sources at random and confirm the sentence attached to each is in the paper.
ai academic research: the retrieval layer that keeps every source checkable
Academic research is the stage where "did the model help?" becomes measurable, because the raw material is a database with an API and the answer is a set of records that either resolves or does not. That is the good news for verification: a source is checkable by construction if it carries an identifier, and the identifier is a command, not a judgement call.
The mechanism behind a usable retrieval layer is worth reading as code, because it shows what "verifiable" means at the level of a single call. A research step sits in front of several backends, each with its own response shape, and its job is to return one uniform hit while recording which backend answered and whether an engine failed. That is the difference between "the search returned nothing" and "the search engine was down" — a distinction a reader of the paper can act on and a silent empty result cannot:
# backend/smartgate/modules/search/algorithm.py — source lines 119–167 (the retrieval dispatch, its fallback and its provenance)
async def search(self):
import time
self.start_time = time.time()
backend = self.settings.resolved_backend()
self.effective_backend = backend
try:
if backend == "firecrawl":
rows = await self._search_firecrawl()
self.result_container.extend("firecrawl", rows or [])
elif backend == "searxng":
rows: list[dict] = []
try:
rows = await self._search_searxng()
except Exception as e:
logger.warning("SearXNG failed: %s", e)
if self.settings.searxng_fallback:
self.result_container.add_unresponsive_engine(
"searxng", str(e)
)
else:
raise
if rows:
self.result_container.extend("searxng", rows)
self.effective_backend = "searxng"
elif self.settings.searxng_fallback:
logger.info(
"SearXNG returned no rows; falling back to DuckDuckGo Lite"
)
try:
fb_rows = await self._search_duckduckgo_lite()
self.result_container.extend("duckduckgo", fb_rows or [])
self.effective_backend = "duckduckgo"
except Exception as fb_e:
logger.warning("DuckDuckGo Lite fallback failed: %s", fb_e)
self.result_container.add_unresponsive_engine(
"duckduckgo", str(fb_e)
)
else:
logger.warning(
"SearXNG returned no rows (searxng_fallback=false)"
)
else:
rows = await self._search_duckduckgo_lite()
self.result_container.extend("duckduckgo", rows or [])
except Exception as e:
logger.warning("Search backend '%s' failed: %s", backend, e)
self.result_container.add_unresponsive_engine(backend, str(e))
return self.result_container
Two properties in that excerpt are the ones to keep. First, the fallback is recorded, not hidden:
when the primary engine raises or returns no rows, the code falls through to a second engine and notes
the unresponsive one rather than returning an empty list, so a later reader can tell a thin result
from a broken source. Second, the backend that actually answered is kept on the result container
— effective_backend is set beside the hits, not reconstructed afterwards from logs. A retrieval step
that names its own source, and says when a source was unavailable, is the smallest unit of the
"checkable by construction" property this page keeps returning to. Everything downstream — the
synthesis, the write-up, the citation list — inherits whatever
discipline this step had.
ai literature review: mapping a field without inventing a citation
This is the stage that gives AI-assisted papers their reputation, and the reputation is earned. The failure is specific and well documented: a model asked for references will produce plausible ones, with authors, years, journals and a title that reads exactly like a real paper, and no such paper exists. Fluency and existence are independent properties, and only one of them is visible on the page.
The evidence is not anecdotal. A 2023 study of a research proposal drafted entirely by a chatbot verified 178 references it had produced: 69 carried no DOI, and 28 could not be found online at all. Roughly one in six references pointed at nothing — the number to hold onto, because it predicts what happens in a draft nobody checks: fabricated citations are scattered through an otherwise real reference list, so spot-checking the first three proves nothing.
The countermeasure is mechanical, and it is the reason this stage is safe to assist. Every reference gets resolved before it enters the list — a DOI, an arXiv ID or a database record that opens and matches the author and year. Retrieval tools such as the Crossref REST API, OpenAlex and the Semantic Scholar Graph API exist precisely to turn a citation string into a checked record, and a lookup that returns nothing is the check doing its job. The cluster page on the AI literature review owns the full method of mapping a field; the rule this page carries is the one-line version: no reference enters the list unresolved, and a model is welcome to draft the list so long as every line is opened.
ai systematic review: the protocol that makes coverage auditable
A systematic review is a literature review with a protocol, and the protocol exists because "I read widely" is not a method. A model changes the review's central risk: not that it finds too little, but that nobody can reconstruct what it searched.
The reporting standard to anchor on is PRISMA 2020, a 27-item checklist built for exactly this problem: which sources were searched, when each was last searched, the full search strategy for each, the eligibility criteria, and how many records were identified, screened, excluded and included. The list is not decoration — it is the difference between a claim about a field and a claim about a search. A paper that reports its query, its databases and its date can be reproduced by a reader with the same access; a paper that reports a synthesis of "the literature" cannot.
A retrieval layer makes this cheaper and easier to fudge. Run the searches through something that returns a date-stamped, uniform result set — the discipline the excerpt above enforces per call, lifted to the whole review. A deep research API is a reasonable home for that: it turns "the model browsed for a while" into a logged sequence of queries with timestamps, the artefact PRISMA item 6 asks for. The failures to watch are familiar — a search strategy described more precisely after the fact than it was run, and an eligibility boundary that moves to include the papers the model happened to find.
ai meta analysis: pooling numbers after a model extracted them
Meta-analysis is the stage where an AI paper can go quietly wrong in a way that survives every read of the prose, because the error is a number in a table, not a sentence. The pooled estimate is the product of many small extractions, and a model that extracts effect sizes, sample sizes and variances from forty papers will make small, plausible mistakes — a confidence interval read as a standard deviation, a p-value read as an effect size, an n taken from the wrong group.
The arithmetic downstream is unforgiving in one direction and forgiving in another. It is unforgiving because a single transposed number can move the pooled estimate past significance, and the summary forest plot will look perfectly ordinary. It is forgiving because the fix is cheap and total: the extraction is a table, and a table can be checked row by row against the source figures. The pattern that works is extract with a model, verify with a person, then compute with code — the model fills the table, a human opens each source and confirms the row, and the pooling is done by a script rather than described by a paragraph. When the model also writes the analysis code, the run is the check: a result a human can reproduce is not a summary of a result, and the difference is the entire protection this page is arguing for.
Two habits close the gap: record the extraction rule before extracting, so which measure to take when a paper reports several is a decision rather than a coin flip discovered later; and keep the extraction table as a supplementary file, because the table, not the forest plot, is what makes the meta-analysis auditable when a reader disagrees.
The three academic accidents, and the check that catches each
The failure modes cluster into three, and each has a specific check that is cheap enough to run on every paper.
| Accident | How it looks on the page | The check that catches it |
|---|---|---|
| A fabricated citation | a real-looking reference; correct format, no such paper | resolve every reference to a DOI or record before it enters the list |
| A wrong statistic | a fluent sentence about a number; no visible error | re-run the test from the extraction table; never quote a computed number |
| A misread figure | a confident sentence about a chart or a table cell | open the figure and read the value; do not let a summary stand in for it |
The three share a property that explains why they recur: each is a claim the reader is least equipped to re-derive at the moment they read it. A citation looks checkable but is rarely checked; a statistic looks settled; a figure is glanced at rather than measured. That is why the three checks above belong in the workflow rather than in a promise: they are the actions a tired author skips, and they cost minutes per paper. A deep research agent that runs the search can be asked to return raw records rather than a synthesis, which is what gives a human something to check the summary against.
Verifiable by construction: the rules that keep a paper checkable
The pattern across the previous sections is one idea: make the paper's claims resolve to something a second party can open. Six rules put that into practice.
- Every reference resolves. A DOI, an arXiv ID or a database record that opens and matches the author and year. No resolution, no reference — the draft is not the list.
- Every number has a run. A statistic in the prose is a copy of a value a script produced; the script and its input table are the artefact, and the sentence merely quotes them.
- Every figure has a source value. A claim about a chart is read from the chart, not from a summary of it, because a summary of a chart is a description of a description.
- Every search has a date and a query. What was searched, where, and when — the PRISMA-shaped record that lets a reader reproduce the coverage rather than trust it.
- Every extraction row is confirmed. A model fills the table; a person opens the source and checks the row; the pooling runs on the confirmed table.
- Every stage has a named owner. The value of the rule is not the signature but the sentence a reader can ask: who checked this, and against what.
None of this is specific to AI, which is the point. The rules are ordinary research hygiene, and the reason they matter more now is that a model makes the production of unverified claims cheap. The defence is not to write more slowly; it is to make each claim arrive already attached to its evidence.
How SmartGate fits
SmartGate is the intelligence layer between an agent and the sources it works with: it runs the retrieval, compression and memory steps of a research pipeline as algorithm primitives, so a workflow does not rebuild them per project. For this page's argument, that matters in one way — those primitives are the retrieval half, the half this page says is safe to delegate.
The five capabilities map onto a paper's early stages rather than its late ones. Research (aggregated search plus fetch) covers the literature search; Context (de-duplication and compression) covers the reading pile; Memory keeps what was found across a longer project; Control puts a budget on the retrieval spend; and Pipeline runs the sequence end to end. None of them writes the statistics or resolves the reference list, which is the division of labour this page argues for: the platform owns the checkable retrieval, the author owns the signable result. A single call that returns the sources, and records which engine answered and when, is the shape the excerpt above shows at the level of one query.
The plan table is a set of operational limits rather than a feature matrix: monthly token caps of 2M, 20M, 100M and 200M+, requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and 2, 10, 30 or unlimited keys per team. If a retention requirement is set by a journal or a funder policy, read it against those numbers before committing to a date; the pricing page is authoritative and carries the current terms. To try the retrieval layer on a real question, start free and run one search before you wire it into anything.
Frequently Asked Questions
Can AI write a whole research paper?
It can produce every part of one, and that is the problem rather than the answer. A model can draft the introduction, method, analysis and discussion, but two of those are only valid if the numbers and sources behind them are real. The working split: let it draft everything, and let a human own the reference list and the statistics, because those are the parts a reader cannot re-derive and therefore the parts that carry the paper's integrity.
How common are fabricated citations?
More common than the fluent prose suggests. In a 2023 audit of a research proposal drafted by a chatbot, 178 references were checked and 28 could not be found online at all, while 69 carried no DOI. Roughly one in six pointed at nothing. Because the invented entries look exactly like the real ones and are scattered through the list, spot-checking a few references does not tell you the rest are sound.
Is it safe to let AI summarise the literature?
Summarising is safe and useful, provided the summary is treated as a reading list rather than a result. The risk is that a synthesis quietly averages away the paper that disagrees, so the summary should keep a pointer from every sentence back to the source it came from, and a human should open a sample of those sources. The summary tells you where to look; the sources are still the evidence.
Should the model run the statistics?
The model may write the code, but the number in the paper should come from a run a human can reproduce. A computed value that a person re-runs is checkable; a value a model simply described is not. Keep the extraction table, run the tests with a script, and quote the script's output rather than a sentence about it.
What is the single cheapest protection?
Resolve every reference before it enters the list. A DOI or database record that opens is a mechanical check, it takes minutes per paper, and it removes the one failure that most reliably turns a finished paper into a correction notice.
Limitations
- This page describes a division of labour, not a guarantee. Keeping the reference list and the statistics human-checked makes those two artefacts reliable; it does not make the framing, the method or the interpretation correct, and those remain the author's to get right.
- The failure statistics are illustrations, not rates for your field. The citation audit quoted above measured one proposal, one model, one time; another discipline, or a model with retrieval attached, will produce different numbers. Treat it as evidence the failure is common, not as a rate to plan against.
- Retrieval layers add a step, not a solution. They make a search reproducible and a source traceable; they do not decide whether the sources found are the right ones, which is a judgement the protocol's eligibility criteria still have to state.
- Vision models improve, and the check stays. Reading charts from a rendered image is getting better, but any figure read this way is a measurement with an error bar, and the cheap protection — open the figure and read the value — does not depend on how good the model is.
- This page quotes one code excerpt, a measured finding rather than a shortcut: see the method note below.
Sources
- Alkaissi H, McFarlane SI. Artificial Hallucinations in ChatGPT: Implications in Scientific Writing. Cureus, 2023; DOI 10.7759/cureus.35179.
- Athaluri SA et al. Exploring the Boundaries of Reality: Investigating the Phenomenon of Artificial Intelligence Hallucination in Scientific Writing Through ChatGPT References. Cureus, 2023; DOI 10.7759/cureus.37432 — the audit of 178 references behind the one-in-six figure.
- Page MJ et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ, 2021; DOI 10.1136/bmj.n71 — the 27-item checklist and the search-reporting items.
- Masry A et al. ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. arXiv:2203.10244 — human-written and generated chart questions, the benchmark behind the "charts are read, not trusted" rule.
- Crossref REST API, OpenAlex and the Semantic Scholar Graph API — the services that resolve a citation string to a record.
- Demand figures in this page are our own measurement: DataForSEO Google Ads, United States, 12-month
window, measured 2026-10-02, recorded in this project's
search_volume.jsonandresearch_brief.md. - Product behaviour and the plan table: read from the product source at the revision pinned in this
project's
pipeline_results.json, read-only, with the plan figures re-verified against the live pricing page on 2026-10-02.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|---|---|---|---|---|---|
| 1 | ai academic research | Search |
backend/smartgate/modules/search/algorithm.py |
119–167 | rule A L2 → slot-proof | 5d88cd6aed7c |
Method note
The matcher pinned 4 of 7 sections for this page (0 abstention(s),
3 no-slice verdict(s)), and one excerpt is quoted above. All four "pinned" sections resolved
to the same symbol — the Search container in backend/smartgate/modules/search/algorithm.py, which
the section keywords reach because they share the words "research" and "search" — so they are a
generic-name collision rather than evidence about writing a paper. Quoting one class four times would
make the page look verified without adding a single fact, so it is quoted once, under the section where
retrieval provenance is the point, and the other six sections are written from the sources listed above.
The product claims were read from the product source at the revision pinned in this project's
pipeline_results.json, read-only, and the plan figures were re-verified against the live pricing page
on 2026-10-02. The quoted block is cut from the slice body and re-asserted against it byte for byte
before publication; no batch fingerprints, auction data or internal hosts appear anywhere in the text.