AI Literature Review: The Systematic Review Pipeline
An AI literature review is the same audit trail a systematic review has always required — a protocol written before the search, a reproducible search string, deduplication, two-stage screening, a PRISMA-style flow count and an evidence table — with a language model doing the mechanical passes and a named human owning every inclusion, exclusion and bias judgement.
Short answer: An AI literature review is the same audit trail a systematic review has always required — a protocol written before the search, a reproducible search string, deduplication, two-stage screening, a PRISMA-style flow count and an evidence table — with a language model doing the mechanical passes and a named human owning every inclusion, exclusion and bias judgement. The tool changes the labour, not the standard.
Key takeaways
- Four artefacts separate a review from a summary: the protocol, the verbatim search string, the PRISMA flow counts, and an evidence table with one row per included study.
- A language model is strongest at the passes that are broad and cheap — synonym expansion, deduplication, abstract triage, field extraction — and weakest at the passes that carry the argument: applying criteria to a borderline record, appraising bias, and verifying a number against its source.
- Reproducibility is the test: a second reviewer who cannot re-run your exact string and land on your exact included set has no way to check your conclusions.
- The next step is to write the protocol before the search and to decide, in writing, which steps a model is allowed to decide and which it may only suggest.
ai literature review: what the artefact has to contain
The phrase covers two very different outputs. One is a search that returns a reading list: a prompt produces twenty plausible papers and a summary of them. The other is a review: an argument whose every claim is traceable to a source, produced by a method that a different team could repeat and get substantially the same included set. The difference between them is not effort or quality of prose. It is the presence of four artefacts, and a page that has all four has a review whether or not a model helped write it.
The protocol is written before the search runs and names the question, the databases, the date window, the languages and the inclusion and exclusion criteria. The search string is recorded verbatim, per database, with the exact date it was executed and the number of records it returned. The flow is a PRISMA-style count: records identified, duplicates removed, records screened, records excluded with reasons, full texts assessed, studies included. The evidence table holds one row per included study and only the fields the synthesis will actually use.
A model can draft any of these in minutes and none of them can be trusted until a person signs the protocol and reads the counts. The tooling question — which product retrieves, which deduplicates, which stores the corpus — is a separate subject, and this cluster's centre page, AI research tool, owns the tool-class taxonomy. This page owns the method: what the review has to contain, and where in that chain a model earns its place.
literature review ai: the division of labour, stage by stage
The useful way to plan an AI-assisted review is not "can the model help" but "at this stage, does a model decide or suggest". Every stage falls on one side, and the borderline stages are the ones that quietly corrupt a review when they are handed over. The table is the whole argument in one place.
| Stage | A model may do | A named human must do |
|---|---|---|
| Protocol and question | Draft a PICO frame; find comparable registered protocols | State the question, register it, own the risk |
| Search string | Expand synonyms; map controlled vocabulary; translate to each database's dialect | Approve the string, the databases and the date limits |
| Deduplication | Normalise identifiers; merge imports; propose duplicate clusters | Confirm which merged records are the same paper |
| Title and abstract screening | Rank and triage; retrieve missing abstracts; suggest an obvious exclude | Decide every borderline record; dual-screen a sample |
| Full-text screening | Pre-extract candidate fields; summarise the methods | Apply the criteria; resolve disagreements; record the exclusion reason |
| Appraisal | Summarise what the paper reports | Assign the risk-of-bias judgement |
| Synthesis | Cluster themes; draft prose from the evidence table | Weigh the evidence and state the uncertainty |
The pattern holds at every row: the model works on the breadth of the corpus and the person owns the decisions that the conclusion rests on. An autonomous retrieval loop — the kind of system that plans several searches and merges the results — is genuinely useful at the top of this table because it widens recall, and its output is a candidate list, not an included set. That retrieval loop is the subject of deep research agent; what this page adds is the receipt the loop has to hand back to the reviewer.
ai for literature review: building the search string you will have to defend
The search string is where an AI-assisted review either becomes reproducible or does not. A defensible string is built in blocks rather than as a sentence: one block for the population, one for the intervention or exposure, one for the outcome or study design, combined with the Boolean operators the database understands. Phrase quotes fix multi-word terms, truncation catches spelling variants, and a controlled vocabulary — MeSH in MEDLINE, Emtree in Embase — catches the papers that never use the author's word. Sensitivity and precision trade against each other: a string that misses a relevant paper fails the review, while a string that returns everything moves the cost to the screening stage.
A model is good at three parts of this and careless about a fourth. It generates synonyms and spelling variants quickly, it translates a finished string from one database's syntax to another, and it will point out a missing block that a tired human skipped. What it also does, reliably enough to matter, is invent index terms that do not exist as controlled vocabulary and silently drop a term when it rewrites the string. So the rule is the one the rest of the method follows: the model proposes, a person freezes the exact text, and the frozen text goes into the protocol with the database, the date and the result count beside it.
That record is what makes the search reproducible, and reproducibility is the only reason the rest of the review can be checked at all. A reader who cannot re-run your string cannot reproduce your set. Where the retrieval is spread across several engines and a deep pipeline, the class of system that automates it is covered under deep research AI; the obligation this section adds is that whatever runs the search must be able to print the string it ran and the number it returned.
ai literature review tool: what the screening layer must do
Whatever the retrieval, the screening layer has a fixed set of obligations, and a tool that cannot meet them will push work back onto a spreadsheet. It must import the standard export formats — RIS, BibTeX, CSV — without losing records. It must deduplicate in two passes, first on a stable identifier and then on a fuzzy title-and-year match, and keep a log of every merge so a mistaken merge can be reversed. It must support screening in two stages, title-and-abstract before full-text, with each record carrying a decision, a reason and the name of the person who made it. It must export the counts the flow diagram needs, and it must let a second reviewer screen blind.
A language model adds value in exactly one place here: triage. It can rank a thousand abstracts by apparent relevance, pre-fill a candidate exclusion reason, and pre-extract the fields the evidence table will want. It should not make the call. The measured failure mode of model screening is not overt error, it is a false-negative rate that moves with the prompt and the model version: relevant records dropped quietly, with no flag on the record that would tell a human to look. The practical guard is to have the model suggest and a person confirm, and to audit a random sample of the model's excludes against human decisions on every review, so the rate is known rather than assumed. Where the screening is wired into a program rather than a product, the integration path is deep research API; the audit obligation is the same either way.
ai systematic review: PRISMA, bias and the evidence table
An AI systematic review is judged by the same reporting standard as a manual one, and that standard is concrete. The PRISMA 2020 statement is a 27-item checklist plus a flow diagram, and the diagram is the part an automation project finds hardest because it depends on counts that must reconcile at four points: records identified by the searches, records remaining after duplicates are removed, records screened, and studies included. The four numbers have to agree with the search log and with the screening decisions; a gap is usually a lost import, not a rounding difference. Registration in a public registry before the search is the second piece of the standard, and it is what separates a review from a retrospective write-up.
The evidence table is where the included studies become analysable. One row per study, and each row carries the design, the sample, the population, the comparison, the outcome, the reported effect and the quality judgement. Two disciplines make it trustworthy. First, every extracted figure is verified against the source at the point of entry, because an error introduced here propagates into every sentence of the synthesis. Second, the quality judgement is a human one. Bias appraisal asks whether the study could have produced the result for reasons other than the effect — randomisation, blinding, missing data, a switched outcome — and that question is answered by reading the methods, not the abstract. A model can pre-fill the table and flag the cells it could not find; assigning the bias rating is the reviewer's work. How a single study is appraised before it earns its row is developed in the next section.
ai research paper: appraising one source before it enters the synthesis
Before a paper is allowed into the evidence table it is read as a methods problem, not as a claim. The reviewer asks what the design was, what it was compared against, whether the primary outcome was specified in advance, whether the sample is the one the conclusion needs, and whether the limitations the authors list are the ones that actually bound the result. The distinction that matters most is between the paper's claim and the data it reports: the abstract asserts, the tables measure, and the reviewer's job is to record what the study measured. A well-written abstract can be honest and still overstate a result the tables do not support, and an extraction sheet filled from abstracts inherits that overstatement.
Model reading of a single paper is genuinely useful for navigation — locating the methods paragraph, the sample size, the stated comparator — and genuinely dangerous when it copies a number out of the abstract straight into a table. The working rule that survives: every extracted figure is checked against the source line, and any quoted phrase is kept only if it is verbatim. The boundary with the sibling page is sharp. Producing a paper — turning a result into a manuscript — is the subject of AI research paper; this section only covers reading one paper as an input to a synthesis, which is a different job with a different failure mode.
research paper summarizer ai: extraction, not summary
A summariser is the weakest link in an AI-assisted review, because fluent prose and faithful prose are not the same thing, and the failure is invisible to the reader. The recurring classes are specific enough to name. A statistic is produced that appears nowhere in the source. A citation is real but does not support the sentence it is attached to. A null or adverse finding is softened into a positive direction, or dropped because it did not fit the narrative. A number is transcribed from the wrong arm of a trial, so an intention-to-treat result becomes a per-protocol one. None of these are caught by a spell-check, and all of them survive into the synthesis if the summariser is trusted as the record.
The correction is to stop asking a model to summarise and to ask it to extract. A typed slot — this paper, this outcome, this measure, this value, this location — is checkable in a way a paragraph is not, and the model should be allowed to return "not reported" rather than forced to fill the field. The provenance stays attached: the raw document is kept and the extracted row points at the table or page it came from, so a second reviewer can re-derive the value without re-reading the whole paper. Spot-checking a sample of extracted rows against their sources turns the error rate into a measured number, which is the only form in which it can be tolerated.
ai researcher: the reviewer of record
Every review names the person accountable for its protocol and its decisions, and no amount of automation changes that. What changes is the shape of the work. The mechanical portion — typing search strings into several interfaces, moving records between formats, formatting a reference list — shrinks, and the judgement portion grows as a share of the day. The skills that carry the review are the ones a model cannot hold: writing a question that can actually be answered, choosing a search that finds the literature without drowning it, resolving a borderline exclusion against a written criterion, reading a methods section for the source of a bias, and stating honestly how much the evidence supports. A reviewer whose only contribution is pressing run on a pipeline has moved the accountability without sharing it.
That is the line this page draws. A model can be a research assistant that widens recall and pre-fills the record; it cannot be the author of record for a claim that no person can trace to a source. The role itself — what a researcher does now, which parts of the craft survive automation and which do not — is the subject of AI researcher. Here it is enough to say that the reviewer of record is a person, and that the protocol says so before the search is run.
What our fetcher hands the review pass, read from our own converter
A review step is only as good as the text it receives, and the conversion from a live page to text is
where that text is decided. Read on 2026-10-07 from
backend/smartgate/modules/fetch/html_converter.py and modules/fetch/models.py.
- Structure is preserved deliberately, not incidentally. The converter is a faithful re-implementation of an upstream HTML-to-Markdown service, and its post-processing is GitHub-flavoured on purpose: tables get their separator rows repaired, code blocks keep a fence and their language, task lists survive. A review pass that can still see a table is a review pass that can quote it.
- The result carries a quality score and a length. Both are cheap to ignore and both decide whether a review is worth doing: a short extraction with a low score is usually a paywall, a redirect or a consent page, and summarising it produces confident nonsense.
- Timeouts are bounded per request, between five and 120 seconds with a 30-second default. For a batch review that means a slow host cannot stall the whole run — it fails one item.
- The honest limit: this step extracts, it does not judge relevance. A literature-review pipeline that skips a relevance filter will faithfully compress an unrelated page, and the compression ratio will look excellent.
Where SmartGate fits
For a team running the mechanical passes on a real corpus, the retrieval and context steps are
where cost and provenance are won or lost, and SmartGate is the layer that runs them under a budget
and leaves a record. smart_search aggregates several engines behind one call so a search strategy
is not re-wired per source; smart_fetch turns a URL or a PDF into markdown so a full text can be
read without a bespoke parser; smart_dedup removes overlapping text across a corpus, which is the
same operation the deduplication stage needs at a coarser grain; smart_context_gate compresses a
long document to a ratio you set before it enters the window; smart_memory holds a team's
extraction vocabulary so a second reviewer inherits the first reviewer's field names; and
smart_budget_guard enforces a hard per-team cap so a screening pass over a large corpus cannot run
away. smart_pipe chains the research, read and remember steps when the passes are routine. These
are five capabilities powered by seven primitives, not seven separate products.
The operational limits are plan properties rather than features: monthly token caps of 2M, 20M, 100M and 200M+, requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and a team holding 2, 10, 30 or unlimited keys. Because the team pays for the platform and shares the saving only past a threshold once its own usage is measured, the cost model suits a review that runs in bursts. Compare the retention figure against your own audit requirement before you commit a date to it — the pricing page is the authoritative table, and it is the one to read rather than a secondary summary.
Frequently Asked Questions
Is an AI literature review acceptable in a peer-reviewed systematic review?
The reporting standard does not ask what tool produced the review; it asks whether the method is transparent and reproducible. A review qualifies when the protocol, the exact search string, the screening decisions and the PRISMA counts are all reported, whoever or whatever ran the searches. Using a model without reporting it, or letting it make the excluded set, is what fails review.
Can a model replace the second reviewer?
No. Dual screening exists to catch the errors a single reader makes, and a model's errors are correlated with the prompt it was given, so it does not provide the independence a second human does. Model pre-screening is useful for triage and for a cheap first pass; the confirmatory screen should still be a person, and a sample should be dual-screened.
How do I stop the model fabricating citations?
By never letting it produce a citation as output. Retrieve the record from a real source, keep the identifier, and let the model only ever select from records that already exist. A fabricated reference is a retrieval failure being papered over, and the fix lives in the retrieval step, not in the prompt.
What has to be in the search string record?
The string exactly as run, the database or engine it ran against, the date it was executed, and the number of records returned. If the string was translated between databases, keep both versions. Without the date and the count, the flow diagram cannot be reconciled and the search cannot be re-run.
How many records is too many to screen?
There is no fixed ceiling, but the ratio is the warning sign: if the search returns tens of thousands of records for a question the evidence base cannot support, the string is too broad and the cost has moved to screening. A tighter string costs recall, so the trade should be made deliberately and recorded, not discovered after the first thousand abstracts.
Limitations
This page describes the method of a defensible review, not a benchmark of any tool or model, and it does not rank products. The division-of-labour table states where judgement is required; it does not claim that every review project has the staffing to dual-screen, and a review run by one person should say so rather than imply a second screener that did not exist.
The product limits quoted above are plan properties, not a comparison, and they change; a compliance or retention decision must read the current pricing table rather than this page. The figures for how often a model drops a relevant record are deliberately not stated as a number here, because they move with the model and the prompt; the honest position is that the rate must be measured per review, not assumed from a vendor claim.
Sources
-
The PRISMA 2020 statement and its flow diagram — prisma-statement.org, the reporting standard this page's four artefacts follow.
-
The Cochrane Handbook for Systematic Reviews of Interventions — training.cochrane.org/handbook, for risk-of-bias appraisal and the review process.
-
PROSPERO, the international prospective register of systematic reviews — crd.york.ac.uk/prospero, the protocol-registration layer.
-
The EQUATOR Network's reporting guidelines — equator-network.org, which collects the standards a review is checked against.
-
The Medical Subject Headings controlled vocabulary — nlm.nih.gov/mesh, the index terms a search string is built from.
-
The keyword figures in this page are this project's own measurement, recorded in
projects/ai-literature-review/search_volume.jsonandresearch_brief.md. -
The section "What our fetcher hands the review pass, read from our own converter" is our own implementation, read on 2026-10-07 from
backend/smartgate/modules/fetch/html_converter.pyandmodules/fetch/models.py(origin/main). It states only what those files state.
Method note
This page carries no code excerpt, and the reading behind that decision is recorded here rather
than left implicit. The slice matcher pinned 3 of 8 sections for this
page (0 abstention(s), 5 no-slice verdict(s)) — but all three pins resolved to the
same generic container symbol, Search, and the section carrying the page's own term was an L0 miss
whose fallback returned unrelated interface helpers. Quoting one generic search container three
times would give the page the appearance of verified code with none of the substance, so every
section above is written from public sources and from the demand this project measured.
The product behaviour and the plan limits were read read-only from the product source at the
revision recorded in this project's pipeline_results.json, with the plan figures checked against
the live pricing page. The keyword figures are this project's own paid measurement, not a
third-party estimate. No code, batch fingerprints, auction data or internal hosts are transcribed,
so there is nothing on this page that has to be asserted verbatim.