SmartGateSmartGate

AI Research Tool Selection: A Buy-Versus-Build Framework

An AI research tool is one of five classes — retrieval, summarisation, literature review, deep-research agent, or a research gateway you build — and the useful decision is which class the job needs, not which vendor has the longest feature list. The head phrase draws about 1,300 searches a month, and almost every page behind it ranks products.

Short answer: An AI research tool is one of five classes — retrieval, summarisation, literature review, deep-research agent, or a research gateway you build — and the useful decision is which class the job needs, not which vendor has the longest feature list. The head phrase draws about 1,300 searches a month, and almost every page behind it ranks products. This page gives the framework instead: what each class actually decides, the eight dimensions that separate candidates inside a class, when buying beats building, and the total cost of ownership you are really signing up for.

Key takeaways

  • Five classes, five questions: retrieval asks can it find the evidence, summarisation can it compress it, literature review can it synthesise many sources, a deep-research agent can it run the loop, and a self-built gateway can we govern the loop.
  • Compare candidates on eight dimensions, of which traceability and data retention are the two that cannot be retrofitted onto a tool after you have chosen it.
  • Buy when the work is commodity and the sources are public; build when the corpus, the credentials or the audit trail is yours alone and no vendor is allowed to hold it.
  • Total cost is subscription plus seats plus integration plus evaluation plus switching — the sticker price is usually the smallest of the five.
  • Start by writing down the one artefact you must be able to produce on demand; every tool decision below follows from that single line.

ai research tool: five classes, and the job each one decides

The market files all of this under one heading, which is why the comparison pages read the same. "Research tool" covers at least five different machines, and most disappointing purchases are class errors rather than vendor errors: someone bought a summariser to do a systematic review, or an autonomous agent to answer a question a single search would have closed.

Class The question it answers Typical shape Failure mode when used for the wrong job
Retrieval Can it find the evidence at all — the right pages, papers and data, from the right sources? Search API, multi-engine aggregation, web fetch, a reading layer over a corpus Confident answers with no traceable source; coverage gaps you cannot see
Summarisation Can it compress one long document into its claims, still attached to the original? Per-document summariser, chat-with-PDF, extraction into structured fields Summaries that read well but drop the caveats and the numbers that mattered
Literature review Can it synthesise many sources under an explicit inclusion rule, and keep them distinct? Review workspace, screening and de-duplication, evidence tables Double-counted studies, silent exclusions, a synthesis you cannot defend
Deep-research agent Can it run the loop — plan, search, read, revise — and stop on a condition you set? Multi-step autonomous research with a plan, a budget and a stop rule Unbounded spend, unverifiable multi-step claims, no reproducible path
Self-built gateway Can you govern the loop — keys, limits, audit, memory — for many callers? Your own research infrastructure behind one authenticated surface Rebuilding research plumbing per team, with cost and access uncontrolled

Retrieval is the class that decides whether the other four can be trusted at all: a synthesis is only as good as the evidence the retrieval step surfaced, and an agent's plan is only as good as the the sources it can reach. That is why the first budget line in this framework is not the model — it is the finding layer, and why a tool that hides its sources is not a research tool in the sense this page uses the term. What a research role does with that evidence, end to end, is a different question and belongs to the cluster's researcher page; this page stays on the buying decision.

The second habit that saves money is naming the unit of work before naming a tool. A single paper, a question, a recurring monitoring job and a literature corpus are four different units, and each one pulls toward a different class. Write the unit down first; the class choice follows almost mechanically.

ai research paper: a paper-shaped task and what it demands of a tool

Papers are the most common unit of research work, and they are hostile to casual tools in three specific ways. First, a paper is a versioned object: the preprint you read in March is not the camera-ready you cite in June, and a tool that cannot tell versions apart will happily let you cite the wrong one. Second, a paper's value often lives in a table or a figure, not in its abstract — a tool that only reads prose returns a summary of the least falsifiable part of the document. Third, a paper has a retraction and correction history that matters more than its content.

A tool that fits this unit therefore has to expose, at minimum: persistent identifiers (DOI, arXiv id), the version it read, the section or page a claim came from, and a way to check whether the work is still standing. Anything less produces the familiar failure — a fluent paragraph with a citation that a reviewer can dismantle in one click. The mechanics of working one paper at a time, from finding it to citing it correctly, are worked through on the ai research paper page; here the point is only that the paper unit is what makes summarisers insufficient on their own, because a summariser compresses text and the paper unit requires provenance.

ai researcher: the workflow a tool has to fit, not replace

A tool that does not fit the existing workflow gets abandoned in a fortnight, whatever its benchmark score. The research workflow is stable and short: frame the question, scan the field, gather and screen sources, read and extract, synthesise, and write with citations. Each step has a natural artefact — a question, a source list, an inclusion decision, an evidence table, a draft — and a tool earns its place only by producing or improving one of those artefacts.

Two consequences follow for selection. The first is that the expensive input is researcher attention, not tokens, so the right metric is not "how much can it generate" but "how much verified reading did it save". The second is that a tool which inserts itself into every step is usually worse than several tools that each own one step cleanly: one that reads well, one that searches well, one that keeps notes, one that writes in the author's own structure. A tool that owns no single step owns nothing, and will be removed the first time it slows someone down.

The workflow itself — what an AI research role actually does and how the pieces hand off — is the subject of the ai researcher page. The selection rule that lives here is narrower: count the artefacts your workflow already produces, and buy only tools that touch one of them.

deep research agent: the autonomy question and where the loop must stop

The phrase "deep research" has been stretched to cover everything from a longer answer to a multi-hour autonomous run, so it is worth being precise — and worth reading what deep research actually is before buying an agent that claims the label. What separates an agent from a one-shot answer is the loop: it plans, executes a step, reads the result, revises the plan, and repeats. That loop is the source of the category's real capability and of its real cost, because each iteration spends tokens and time and can wander.

Three design properties decide whether the loop is useful rather than merely long. It needs a plan that is inspectable data, so a human can see where it went and why. It needs a stop condition that is not "the model said it was done" — a maximum number of steps, a token or currency budget, a wall-clock limit. And it needs provenance per step, so that the paragraph it finally produces traces back to specific retrieved spans rather than to the loop's own summary of itself. A deep-research agent without a stop condition is a runaway bill; without step provenance it is a better-written guess.

Where the category ends and how the loop is actually built — planning, tool choice, reflection, stopping — belong to the deep research ai and deep research agent pages. The commissioned run itself — what to brief, what to accept, what to keep as a reusable asset — is the subject of ai deep research, and the open-source implementation layer, including how to audit a repository before self-hosting it, is the subject of deep research on GitHub. The selection question that stays on this page is whether the job deserves a loop at all: if a single retrieval and a read would answer it, an agent is an expensive way to make a simple task slower and less legible.

deep research api: programmatic access and the integration tax

The moment research stops being a person typing and becomes a pipeline — nightly monitoring, an agent that calls research as a tool, a product feature that summarises a topic on demand — the unit of purchase changes. You are no longer buying a workspace; you are buying an API surface, and the questions that decide the choice are operational rather than editorial.

Check the following before an API proves itself in staging rather than in production. Rate limits and quota shape: per-minute requests per key, monthly caps, and what happens when you cross them — a clean error, a queue, or a silent slowdown. Pagination and depth: whether a query returns a bounded page you can page through, or a fixed result set you cannot extend. Determinism: whether the same call twice returns the same evidence set, which is what makes an automated result reviewable later. Provenance in the response: the source URL and span behind each returned claim, not just the claim. Retry and idempotency semantics: whether a retried call bills twice or duplicates a downstream write. Cost per successful call, measured on your own traffic, because the advertised price is per request and your request is not the average one.

These are the same properties any backend team already asks of a third-party dependency, applied to a research call. The endpoint shapes, pagination and error contracts that matter most are collected on the deep research api page; on this page they enter the framework as integration cost — the line item teams most often forget when they compare two subscriptions on price alone.

ai research agent: the eight dimensions that separate candidates

Once the class is right, the shortlist inside it is decided by a small number of properties that do not depend on branding. Eight of them, in the order a serious evaluation tends to test them:

The same eight apply when the candidate is an ai research agent rather than a single-purpose tool; what changes is which of the eight carries the risk.

  1. Coverage and source diversity. Which sources can it reach — open web, specific journals, internal corpora — and does it tell you which it used? Coverage you cannot see is coverage you cannot trust.
  2. Traceability. Does every claim carry a link to the exact span it came from? If the tool returns prose and a bibliography but not the mapping between them, it is a writer, not a research tool.
  3. Reproducibility. Run the same query twice; do you get the same evidence set, or a different one each time? Reproducible evidence is what lets a second person check the first person's work.
  4. Freshness and change tracking. Does it know when a source it cited has changed or been retracted, or is the citation frozen at the moment of capture?
  5. Context handling. How does it treat a long source — truncate, summarise, or retrieve the relevant span — and can you see which of those it did?
  6. Retention and privacy. What happens to the documents and queries you send; where are they stored; for how long; can you delete them? For regulated or confidential material this dimension is a gate, not a preference.
  7. Limit and cost predictability. Per-key rate limits, monthly caps and overflow behaviour, so that a spike is a number you can plan for rather than an incident.
  8. Integration surface. A documented API or an MCP-style tool interface, structured export, and the ability to script the whole thing rather than click it.

Two of these dominate. Traceability and retention cannot be added later: a tool that does not return source spans on day one will not return them after a feature request, and a tool that stores your corpus in a place you cannot audit has already taken a decision you were supposed to make. Rank on those two first, and let the other six break ties.

ai research platform: buy versus build, and the trigger that decides it

The buy-versus-build question is not philosophical; it turns on four concrete triggers. Build when the corpus is yours alone (proprietary data, licensed content, internal documents that must not leave), when the credentials and the write path must stay inside your boundary, when the volume is high enough that per-seat pricing is no longer rational, or when the audit trail is a regulatory requirement rather than a nice-to-have. Buy when the opposite holds: public sources, commodity tasks, low volume, and no external obligation on the record.

That trigger has an operational shape. A self-built research gateway is not one component; it is the same five capabilities every serious platform has to provide, and a team that builds without naming them rebuilds them ad hoc:

Capability What it has to do What "buy" gives you instead
Research Search and fetch across the sources you allow, with the source attached to the result A vendor's coverage, bound to their connector list
Context Compress and de-duplicate long material before it reaches the model A per-document summariser with no cross-document de-duplication
Memory Persist what a team has already established, with retention you control Usually the vendor's store, on the vendor's terms
Control Enforce per-key limits and a per-team budget at the moment of the call A spend report that arrives after the invoice
Pipeline Orchestrate research, read and remember into one callable job Point tools you glue together yourself

The honest reading of the table is that "build" is not one project but five, and that most teams that think they are choosing between buying and building are really choosing how many of the five they will own. Owning all five is a platform commitment; owning none is the fastest route to value but the slowest to a defensible audit trail.

ai literature review tool: the synthesis class, kept separate on purpose

Literature review deserves its own class because it is the one research job where counting and exclusion matter more than fluent prose. A review is a claim about a body of work: how many studies met the inclusion rule, how many were excluded and why, how the remaining ones were combined. Get one of those wrong and the review is not merely weak — it is not a review.

The tools that fit this class share three features a generic summariser lacks: an explicit inclusion/exclusion rule applied and recorded, de-duplication across databases (the same paper indexed three ways must count once), and an evidence table that keeps each source separate while the synthesis is written. Without de-duplication a "synthesis of 40 sources" is routinely a synthesis of 27, and the number in the draft is wrong in a way no reviewer can catch from the prose. The discipline of running that process, from search string to screened set to written synthesis, is the subject of the ai literature review page. Here it enters the framework as a boundary: if your job is a review, a per-document summariser is the wrong class no matter how good it is.

Total cost of ownership: the five line items, not the sticker price

Comparing two research tools on subscription price is the most common way to buy the more expensive one. Five line items make up the real number, and only the first is visible on the pricing page.

Line item What it covers Why it is easy to miss
Subscription and seats The licence, per user, per month Seat creep as a team grows past the plan you budgeted
Integration Gluing the tool into your stack: auth, export, retries, the API surface Weeks of engineering that never appear on an invoice
Evaluation Building the small case set that proves the tool still works after a change The work that turns a subscription into a reliable dependency
Retention and compliance Storage, deletion, and the audit evidence an auditor or a client will ask for Usually charged as an enterprise add-on, or not offered at all
Switching The cost of leaving: re-indexed corpora, rewritten automations, retrained habits The number nobody models until they need it

The three that decide most outcomes are evaluation, retention and switching. A tool you cannot evaluate is a tool you cannot defend when its output regresses; a tool whose retention you do not control is a liability the moment the material is confidential; and a tool with a high switching cost is one you will keep using long after it stopped fitting, which is the most expensive behaviour of all. When two candidates are close on the first line item, the cheaper one is usually the one with the lower switching cost, not the lower price.

How the results are ranked, read from our own container

Every research tool ranks something; the question is whether the ranking is a mechanism you can reason about. Ours is, and it is short enough to quote in full: read on 2026-10-07 from backend/smartgate/modules/search/result_container.py.

  • The rank is built from positions, not from a model. A result accumulates the positions it appeared at, and the score is computed from those positions — which means a page that several engines put near the top ranks high, and a page one engine returned twice does not count twice. Agreement is the signal.
  • The container is adapted from a search-engine implementation, not invented here (the SearXNG result container, with its metrics and web-framework dependencies removed). The scoring function itself is copied unchanged, so the ordering behaviour is the upstream one rather than a house heuristic.
  • One honest gap, stated plainly. In this copy the per-engine weight lookup was simplified to a neutral value, so the ordering today comes from positions and the priority setting rather than from per-engine weights. If you are comparing our results to another multi-engine aggregator, that is the difference you will see.
  • The practical consequence for a research tool: "multi-source" is a measurable property of your ranking, not a claim about your provider list. Count the engines behind each answer, or you are taking the word of the tool that asked them.

Where SmartGate fits

SmartGate is a research infrastructure layer, not one of the classes above. It is not an assistant and not another research application: it is the place where the research, context, memory, control and pipeline capabilities a self-built gateway has to provide already exist behind one authenticated surface, so a team can buy its way out of building all five without giving up the record.

The mechanism is five capabilities powered by seven algorithm primitives. Research is smart_search plus smart_fetch — multi-engine search and a fetch that returns clean markdown. Context is smart_dedup plus smart_context_gate — cross-source de-duplication and configurable compression before text reaches the model. Memory is smart_memory, a team-level store so the same finding is not rediscovered every run. Control is smart_budget_guard, with check, count and record against a hard ceiling. Pipeline is smart_pipe, which chains research, read and remember into one callable job — the piece that turns separate primitives into a workflow an agent, or a person, can call once. That split maps onto the buy-versus-build table above: it is the "buy" column, filled in.

The plan table is about operational limits, not features — monthly token caps, requests per minute per key, audit-log retention and team-key counts all move with the tier, and the pricing page is the authoritative table. Read those numbers against the retention line item in the total-cost model above, because audit-log retention is the one column that answers to a compliance deadline rather than to a preference. Billing is deliberately aligned with the saving: you pay for the platform, and a share is taken only once the platform has saved enough to clear a floor — a model that makes the platform's cost self-funding rather than a flat rent.

How to get started

The first four steps below need no purchase at all.

  1. Write the one artefact you must be able to produce on demand — a cited memo, a screened source set, a nightly brief — and the unit of work behind it. That sentence removes most of the wrong classes immediately.
  2. Classify the job using the five-class table, then shortlist inside the one class it names, scoring candidates on the eight dimensions with traceability and retention first.
  3. Run the four buy-versus-build triggers against the shortlist. If none fires, buy; if any fires, write down which of the five platform capabilities you are taking on.
  4. Price the decision with all five total-cost line items, not the subscription, and give a weight to switching cost before you commit a team to a workflow.
  5. If the answer is "buy the infrastructure layer", the fastest test is to connect a client to the research primitives and run one real query end to end — start free with an MCP-speaking client, then read the pricing page once your real call volume tells you which tier you actually need.

Frequently Asked Questions

Is an AI research tool the same thing as a chatbot with web search?

No. A chatbot with search is a single retrieval plus a single answer; a research tool has to keep the evidence attached to the claim and make the same query reproducible. The distinction is traceability, not fluency. A tool that returns confident prose with no link from sentence to source is answering, not researching.

Do we need a deep-research agent, or is search enough?

Most questions are closed by one retrieval and a careful read, in which case an agent adds cost and opacity for nothing. An agent earns its place when the answer genuinely requires several dependent steps that cannot be known in advance, and when you can give it a stop condition and a budget. If you cannot state the stop condition, you are not ready for the agent.

When should we build our own research platform instead of buying?

Build when the corpus, the credentials or the audit trail has to stay inside your boundary, when volume makes per-seat pricing irrational, or when a regulator or client requires the record. Short of those triggers, buying the infrastructure is faster and owning it is rarely the differentiator the team imagines it to be.

How do we compare two tools without a benchmark we trust?

Score both on the eight dimensions, with traceability and retention first, then run one real task through each and check three things: whether every claim traces to a source span, whether the same query returns the same evidence, and what it costs on your own traffic. A benchmark you did not run tells you how the tool behaves on someone else's work.

What does audit-log retention have to do with choosing a research tool?

For any regulated or client-bound material, the record of who queried what, and for how long it is kept, is part of the deliverable rather than an operational detail. If the retention window is fixed by an external deadline, compare it against your obligation before you commit a date to anyone — and pick the tool whose retention you can actually configure.

Limitations

This page is a decision framework, not a benchmark: no product is ranked and no score is claimed, because the right answer depends on where a team's corpus, credentials and record already live. The five-class taxonomy is a selection aid, not a claim that any named product belongs cleanly to one class — real tools overlap, and a platform can appear in several rows at once.

The demand figures quoted above are our own measurements for the United States over a twelve-month window and are recorded in this project's measurement files; they describe how many people search, not how much the topic is worth to any particular team, and they will age. The plan limits cited in the pricing discussion are operational numbers that change with the plan and should be read from the current pricing page rather than from this page.

Because this page carries no code excerpt — the reason is recorded in the Method note below — it makes no line-numbered or implementation-level claim about any tool, including ours. The eight evaluation dimensions describe what to check; they are not a promise that any specific tool passes them.

Sources

  • OpenAI's deep research announcement — Introducing deep research, for the multi-step research-agent shape and its stop-and-budget concerns.

  • Google's deep research rollout — Gemini Deep Research, as the consumer-facing form of the same category.

  • Anthropic's engineering note — Building effective agents, for the workflow-versus-agent split and the role of an explicit stop condition.

  • Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020) — arxiv.org/abs/2005.11401, the retrieval-plus-generation pattern behind the retrieval class.

  • Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022) — arxiv.org/abs/2210.03629, the interleaved reason-and-act loop a research agent runs.

  • NIST — AI Risk Management Framework, for the record-keeping and traceability expectations behind the retention dimension.

  • The Model Context Protocol specification — modelcontextprotocol.io/specification, for the tool-call interface an integration surface is expected to speak.

  • Product landing pages that dominate the head SERP, read as examples of the class rather than as recommendations: Consensus · Elicit · Scite.

  • Demand figures come from our own paid measurement run and are recorded in this project's search_volume.json and research_brief.md; the product behaviour and plan limits above were read from the product source at the pinned revision and the plan figures re-checked against the live pricing page on 2026-10-02.

  • The section "How the results are ranked, read from our own container" is our own implementation, read on 2026-10-07 from backend/smartgate/modules/search/result_container.py (origin/main). It states only what those files state.

Method note

This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned 7 of 8 sections for this page (0 abstention(s), 1 no-slice verdict(s)), but all seven pins resolved to one generic symbol — a Search container — because the matcher's local rule accepts a symbol whose name is a substring of the keyword, and every phrase in this lane contains "research". Seven sections sharing one generic container is not evidence about any one of them, so the page is written from sources instead: a pinned generic would have given it the shape of a verified article with none of the substance. Product claims were read read-only from the product source at the revision the slice run recorded, and the plan figures were re-checked against the live pricing page on 2026-10-02; the demand figures are this project's own measurement. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.

The slice run for this page recorded 7 of 8 sections pinned, 0 abstention(s) and 1 no-slice verdict(s); BLOCKS is empty because those pins collapse to a single generic symbol rather than to section-specific evidence, as the Method note above explains.