SmartGateSmartGate

Deep Research on GitHub: Open-Source Implementations

Deep research on GitHub is four kinds of repository — a full autonomous agent, a minimal reference implementation, a vendor-alternative app, and a framework example — all running one loop: plan sub-questions, search and scrape, synthesise, cite. Finding a repo is easy; vetting one is the skill.

Short answer: Deep research on GitHub is four kinds of repository — a full autonomous agent, a minimal reference implementation, a vendor-alternative app, and a framework example — all running one loop: plan sub-questions, search and scrape, synthesise, cite. Finding a repo is easy; vetting one is the skill.

Key takeaways

  • The open-source deep research category is small and readable: four repo families cover almost everything you will find when you search GitHub, and three of them are under a few thousand lines.
  • A repository is not a product. What a repo leaves out — the stop condition, the storage layer, the licence boundary — is what you inherit the moment you run it on your own keys.
  • The head term is dominated by listicles and star counts; the useful test is not popularity but whether the loop can be stopped, inspected and made reproducible on your own sources.
  • Free and open here means no licence fee, not zero cost: the repo is free, but every run spends model tokens and search-API calls that you own.
  • Clone one minimal repo this week and run a single query end to end before you compare any two of them — one real run tells you more than a dozen README files.

deep research github: the four repo families you will find

Searching GitHub for deep research returns roughly the same four shapes, and knowing which shape a repository is tells you faster than its README whether it can be trusted with a real question. The first family is the full autonomous research agent: it takes a question, plans sub-questions, searches, scrapes, writes a cited report. GPT Researcher is the canonical example — it describes itself as an autonomous agent that conducts deep research on any data using any LLM provider, and has grown a recursive "deep research" mode with tree-like exploration plus optional MCP integration, so a run can also read a specialised data source alongside the web. Whether that class of tool is even the right one for a job is the first question the AI research tool selection framework answers.

The second family is the minimal reference implementation: a small, readable agent whose whole purpose is to show the loop. dzhng/deep-research is the best known — an MIT-licensed repository that states its goal is the simplest implementation of a deep research agent and deliberately keeps the project under five hundred lines so it is easy to understand and build on. If you want to know what the category is, read this family first; a repository you can read in one sitting is a specification you can also audit.

The third family is the vendor-alternative app: a clone of a hosted product's experience, usually with a web interface. open-deep-research is explicitly positioned as an open-source alternative to a hosted deep research feature and ships a multi-model selector, which makes it the family to look at when the goal is a self-hosted experience rather than a library. The fourth family is the framework example: a model lab or framework publishing a reference agent built on its own scaffolding, such as the open deep research example that Hugging Face shipped with smolagents.

Adopting the wrong family is the most common first mistake: a library gives you a function to embed, an app a UI and an opinion, an example the framework's idioms, a minimal repo the mechanism with nothing else. Before any of that, read what deep research means as a category — the label covers both a two-minute multi-step search and a multi-hour autonomous run, and the repository you clone will only do the one it was built for.

perplexity deep research free: closed tiers and the repos that replace them

A large share of the demand behind this topic is people looking for the capability without a subscription: a hosted deep research feature with a free allowance, or an open repository that does the same job on their own keys. The hosted side is real but constrained — free access to a vendor's deep research mode is a rate-limited taste, gated behind an account and a product roadmap you do not control. The host decides how many runs, which model, how deep the loop goes, and when the allowance resets; none of that is a criticism, it is simply the trade you accept for not running anything.

The open repositories are the other half of the answer: a readable minimal agent, a permissive licence and no per-seat cost are all available, and what you give up is the hosted tier's polish. "Free" moves the cost from a subscription line to an engineering line — paid in the time to stand the stack up and the tokens each run spends.

That trade is reasonable for a team with engineering capacity and unreasonable for someone who wants an answer this afternoon. What is the same across the hosted and open question is the family map — a full agent, a minimal implementation, a self-hosted app, a framework example — and the decision is really about which cost you would rather carry; this page stays on the open-source side of it.

deep research ai free: what free actually costs to run

The phrase "free" is doing two jobs here and it is worth separating them. Free as in licence: the code is published under a permissive licence — the minimal dzhng/deep-research repository is MIT, as is open-deep-research — so you can read it, fork it, modify it and ship it without a fee. Free as in cost to operate: every run of any of these repositories spends money you supply, because it calls a language model and one or more search and scrape APIs. The repository itself is the only part that is genuinely free.

That operating cost is not decoration; it is the number that decides whether a run is worth doing. The GPT Researcher documentation reports a deep research run costing around forty cents using a high-reasoning model at high effort, and roughly five minutes per run, with the note that the figure moves with the model you choose. Read those two numbers together and the economics of the category become clear: a few cents on a cheap model is a casual experiment, tens of cents on a strong model is a report you budget for, and either way the spend is per run and belongs to you. A local or small model can collapse the token line, but it usually trades away the faithfulness the report depends on.

The practical consequence is that "free" reads as "no licence fee, plus a metered run": decide what one run may cost and how long it may take, because the repository will not decide it for you. That is the discipline a hosted product enforces behind its own meter.

deep research llm: choosing the model the repo calls

Every repository in this lane is a thin orchestration layer over one or more language models, so the model choice is not a footnote — it is most of the result. The mature repos are deliberately model-agnostic: GPT Researcher advertises any LLM provider, and open-deep-research exposes Google, OpenAI and Anthropic models side by side in its configuration, with local models as an option. The minimal repos are the opposite and pick a default for you, which is convenient and also a decision you have inherited rather than made.

When you choose, rank the candidates on faithfulness to the retrieved sources, not on general reasoning benchmarks. A research report's failure mode is a fluent sentence that its own citation does not support, so the model that matters is the one that stays closest to the spans it was handed. The GPT Researcher configuration documentation makes exactly this point and publishes a small table of hallucination rate and factual-consistency figures per model — with the honest caveat that the measure is faithfulness, not quality, and that a weak model can score well partly because it adds less. Use a table like that to narrow the field, then compare two models on your own queries before you commit.

Three wiring details matter as much as the model name. First, where the model is called from: a hosted endpoint is simple and ships your prompts off-site, a local model keeps them in your boundary and costs you hardware. Second, how the repo caches and truncates context, because a loop that re-reads a growing context pays for it on every step. Third, whether the repo lets you pin the model and the temperature: a report you cannot reproduce is a report you cannot defend, and reproducibility starts with the model call. What the loop does with that context across many steps is the mechanism covered by how a deep research agent runs its loop; the point here is the narrower one, that the model is a configuration decision you own.

deep research open source: licences, forks and what open buys

Two things travel with an open repository and both are easy to under-read: the licence and the fork graph. On the licence, the split in this category is between permissive and copyleft. The minimal implementations tend to be MIT — permissive, short, easy to comply with — while larger agents have been published under Apache-2.0, which adds an explicit patent grant and a notice requirement. Neither is a trap, but they answer different questions: MIT is the fewest strings, Apache-2.0 is the version a legal team is happiest to see for a patent-sensitive use. Read the actual licence file in the repository you adopt; the label in the sidebar is not always the whole story.

The fork graph is the second signal. A healthy repository in this lane shows active forks and ports — the minimal agent, for instance, has attracted a Python port because its authors kept the code small enough to reimplement. Forks tell you the mechanism is legible and worth carrying to another stack; they also warn you that the thing you found may itself be a fork, a layer behind the upstream, or a copy that never tracked a security fix. When two repositories look identical, the useful question is which one is upstream and which has a maintainer.

What open buys you is concrete: to read what the agent does with your sources, run it inside your own network, change the retrieval step, avoid a per-seat fee. What it does not buy is a guarantee — no SLA, no roadmap you influence as a customer, no one whose job is to notice when an upstream change breaks your fork. Open source moves the trust decision from a vendor's terms of service to your own review, which is a better place for it provided you actually do the review — the same AI research agent role a team owns when it runs the workload end to end rather than renting it.

gemini deep research free: where the closed tier stops

The hosted deep research features are the reason "free" keeps appearing in these queries, and the honest description of a free tier is a boundary rather than a product. A vendor's free deep research access is typically a limited number of runs per period on a fixed model with fixed depth, available in some regions and not others. That is enough to evaluate the experience and not enough to run a workflow, and the moment the workflow matters you are back to a metered plan or to the open repositories.

The open alternative for exactly this case is a repository like open-deep-research, which is positioned openly as an alternative to a hosted deep research feature and lets the operator select the model and the search backend. The trade is legible: you lose the vendor's polish, region handling and managed rate limits, and you gain control of the model, the sources and the cost. For a solo researcher the hosted free tier may genuinely be enough; for a team that needs to run the same research on its own corpus every week, the closed tier's ceiling is the reason the repository exists.

Where the closed tier stops is therefore not a failure but a seam. Everything on the hosted side is optimised for a person typing a question; everything on the repository side is optimised for a system calling research as a step. If your use is a person, the free tier is often the right answer and a repo is over-engineering. If your use is a pipeline — a nightly brief, an agent that researches as a tool — then the boundary is exactly where a repository, and eventually a governed endpoint, takes over. The programmatic side of that handover is the subject of the deep research API surface.

gemini deep research tool: the frameworks that wrap the loop

The fourth family is worth its own section because it changes what you are adopting. A framework example — Hugging Face's open deep research built on smolagents is the reference case — gives you a deep research agent written in a particular scaffolding's idioms, with the framework's abstractions, conventions and upgrade path woven through it. The same is true of the LangChain-family examples, which ship a research agent as a graph of steps. Adopting one means adopting the framework: you inherit its model of tools, its state handling and its release cadence along with the research loop.

That is a benefit when you are already in the framework — the agent speaks abstractions you use and the maintenance is shared — and a cost when you are not, since you take a large dependency to get a loop you could have read as five hundred lines of standalone code. The choice is less about the research logic, nearly identical in both, and more about how much scaffolding you want to own.

The practical test is whether the framework earns its weight for the rest of your system. If you have other agents, tool calls and state to manage, the framework is buying you coherence and the example is close to free. If deep research is a single standalone step, the minimal repo keeps the surface small. A middle path that has become common is MCP: a repository that exposes its research agent over the Model Context Protocol can be called from a host application without either side adopting the other's framework, which is how GPT Researcher's separate MCP server is meant to be used. The standalone repositories and the framework examples are converging on that interface, which is worth knowing before you tie yourself to one scaffolding's graph.

ai deep research free: a first run you can actually trust

The fastest way to evaluate any of these repositories is to run exactly one query through one of them and watch four things. First, what it retrieves: run a question whose answer you already know and watch the domains and pages it visits — a repo that searches two engines and dedupes is visibly different from one that punches a single query into a single API. Second, how it stops: find the loop's exit and confirm it is a counter you set, not the model's own claim that it has finished. A run that only ends when the model says so is a run that ends when your budget does.

Third, where it stores: identify what the run writes to disk, to a database, or to a vector store, and where, because "free and open" says nothing about whether your queries and sources leave the machine. Fourth, how it cites: open the finished report and check that a sentence traces to a retrieved span, not to the model's memory of the topic. A report where the links resolve but do not say what the sentence claims is the signature failure of this whole category, and one real run exposes it faster than any README.

Do that once and you will know more about the repository than its star count tells you: whether it fits your sources, whether you can stop it, whether its storage suits your rules, and whether its output is auditable. Only then is comparing two repositories meaningful. This is also the point at which the review-shaped work separates from the loop-shaped work: running many sources under an explicit inclusion rule is a discipline of its own, covered by the AI literature review workflow, and a general deep research agent is not automatically good at it.

How to vet a deep research repo before you clone it

Four questions, in order, decide whether a repository is worth an afternoon. They are deliberately the same four a security reviewer would ask, because adopting an agent is adopting a network-facing program that spends your money and handles your data.

Question What to look for The failure it catches
What does it retrieve? Named, replaceable search and fetch backends; de-duplication across engines A repo hard-wired to one API you cannot swap or self-host
How does it stop? An explicit step, token, time or novelty cap enforced in code between steps A loop that runs until the model says it is done, i.e. until you notice the bill
What does it keep? A documented storage layer, local by default, with retention you can set Sources and queries quietly shipped to a third party store
What does the licence allow? The actual licence file, and whether your intended use complies A copyleft surprise discovered after a fork has shipped

The order matters: retrieval and stopping define what a repository is — change either and you have a different agent — so both must be replaceable and visible, while storage and licence are cheapest to check now and most expensive to discover later. A repository whose code answers all four, even when the README does not, is the one to keep.

The passes a search result takes before a research run reads it, read from our own module

Every claim above about what a repository retrieves has a counterpart in our own stack, and the counterpart is a pipeline rather than a single API call. One search in our gateway lands on backend/smartgate/modules/search/algorithm.py and modules/search/result_container.py, read on 2026-10-08, and leaves only after a fixed sequence of passes.

  1. A backend is chosen once. Search.search() resolves exactly one of Firecrawl, a self-hosted SearXNG or DuckDuckGo Lite from the provider setting: auto prefers the Firecrawl Cloud key, then the SearXNG URL, then DuckDuckGo. A forced provider whose credential is missing falls back to DuckDuckGo rather than failing, so the engine that actually ran is a recorded outcome and not the setting that was typed.
  2. Every hit is normalised to four fields. _normalize_hit reduces each provider's row to title, url, content and engine, the snippet being the first non-empty of markdown, description, snippet or content, so no later step has to know which engine answered.
  3. The backend caps the page. Firecrawl slices its response to max_results; SearXNG fetches in pages of twenty, merges them and slices to the requested count; DuckDuckGo Lite stops as soon as it has the rows it needs.
  4. Non-results are routed out of the list. ResultContainer.extend sends a hit carrying suggestion, answer, correction, infobox or number_of_results to its own bucket, so the main result list holds only real hits.
  5. Duplicates collapse, then the list is scored. Main results merge by a hash of url plus title: the same page found by a second engine becomes one entry, keeps the longer content and title, and accumulates its positions. close() then scores each entry with SearXNG's own calculate_score — engine weight is flattened to 1.0 in our copy — and get_ordered_results() sorts the list by score.

That is the answer to "how many passes does a search take": the engines that answered, a normalisation, a per-backend truncation, a routing pass and a merge-and-score pass stand between a query and the context a research step is allowed to read.

Where SmartGate fits

SmartGate is not one of the repositories above and does not try to be: it is the layer a self-hosted deep research stack runs on when the loop has become something a team depends on. The reason to put a gateway between an open-source agent and the outside world is the same reason the vetting checklist exists — every run needs a record, a ceiling and a guard, and those three want one place to live rather than one per fork.

Concretely, the agent's calls become governed primitives. Research is smart_search and smart_fetch, so the retrieval step a repository left open is a call with an audit row instead of a bare HTTP request. Context is smart_dedup and smart_context_gate, which matter because a repo that re-reads a growing context pays for it on every loop iteration and can crowd out the evidence that contradicts its own draft. Memory is smart_memory, so the same finding is not rediscovered every run. Control is smart_budget_guard, which turns the stop condition from a line of code in the agent into a budget check the platform enforces. And smart_pipe orders the research, read and remember steps into one callable job, so the parts of the loop that are identical across every repository are not re-implemented per fork.

The plan limits are operational, not features: monthly token caps of 2M, 20M, 100M and 200M+, requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and a team holding 2, 10, 30 or 9999 keys. Those numbers sit directly against the two questions the checklist puts first — how a run stops and what it keeps — and the authoritative table is on the pricing page; read it there rather than from this sentence. Billing is aligned with the saving: you pay for the platform, and a share is taken only once the platform has saved enough to clear a floor.

How to get started

The first steps cost nothing but an afternoon, and none of them requires a purchase.

  1. Pick the family before the repository: a library if you are embedding the loop, an app if you want a self-hosted experience, an example if you are already on that framework, a minimal repo if you want to read the mechanism.
  2. Run one query through one repository on your own keys, then apply the four-question checklist to what you saw rather than to what the README claims.
  3. Fix the two things an open repository leaves to you: point the loop at a stop condition you control and decide where its sources and queries are allowed to be stored.
  4. Move the agent's search, fetch and memory calls behind one governed surface once more than one person or job depends on a run — start free with an MCP-speaking client, and read the pricing page once your real call volume tells you which tier you need.

Frequently Asked Questions

Is there a deep research agent on GitHub I can use for free?

Yes. The code is genuinely free under a permissive licence — a minimal MIT-licensed agent, a larger autonomous agent and a self-hosted app are all public. What is not free is running one: every run spends model tokens and search-API calls you supply, so plan the per-run cost before you clone anything.

How do I know an open-source research agent will not hallucinate?

You do not, and neither does the repository. Every agent in this category drafts from retrieved spans and can drift from them. The defence is mechanical rather than a promise: require that each sentence trace to a span, keep the checking step separate from the writing step, and read the citations of one real run before you trust the next.

Should I fork a minimal repo or adopt a full agent?

Fork the minimal one when you want to understand or modify the loop, because a few hundred readable lines are a specification you can audit. Adopt the full agent when you want citations, de-duplication and a report pipeline already built.

What is the licence situation for these projects?

It varies by repository and it matters — the minimal agents tend to be MIT and some of the larger ones are Apache-2.0, which adds an explicit patent grant and a notice requirement. Read the licence file in the repository you adopt rather than trusting the sidebar label, especially if you intend to ship a derivative.

Does running my own agent mean my research data stays local?

Only if you make it so. A local model keeps prompts on your machine, but the default search and scrape backends still send your queries to third-party APIs, and the default storage may not be local either. Decide explicitly which parts of a run are allowed to leave your boundary before you send it a sensitive question.

Limitations

This page surveys repositories, not products: it names families and a handful of public examples, and it ranks none of them. Star counts, licences and defaults change, so every claim here should be checked against the repository as it stands today rather than read as a permanent description. The four-question checklist is a triage aid, not a security audit; a repository that passes it can still have a dependency you have not read.

The page cannot tell you whether deep research is the right tool for a given question. Most questions are answered by one retrieval and a careful read, and an agent adds cost and opacity to that for no benefit. The loop earns its place only when the answer needs several dependent steps that cannot be known in advance and you can state how the run should stop — if you cannot state the stop condition, you are not ready for the agent, open source or otherwise.

Finally, this page carries no code excerpt and the Method note below records why. What follows from that honestly is that it makes no line-numbered or implementation-level claim about any repository: the families, the checklist and the sources are the whole of what is offered.

Sources

  • GPT Researcher — github.com/assafelovic/gpt-researcher, for the full autonomous-agent family, its recursive deep research mode, its stated cost and time per run, and its optional MCP integration.
  • The minimal reference agent — github.com/dzhng/deep-research, an MIT-licensed repository that keeps the implementation deliberately small so it can be read and reimplemented.
  • The self-hosted alternative app — github.com/btahir/open-deep-research, positioned as an open-source alternative to a hosted deep research feature with a multi-model selector.
  • Hugging Face — Open-source DeepResearch: Freeing our search agents, for the framework-example family built on smolagents and the crop of community implementations it names.
  • Anthropic — Building effective agents, for the workflow-versus-agent distinction and why an explicit stop condition is what keeps an open loop useful.
  • Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022) — arxiv.org/abs/2210.03629, the interleaved reason-and-act loop every repository in this lane implements.
  • Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020) — arxiv.org/abs/2005.11401, the retrieval-plus-generation pattern the retrieval step of these agents formalises.
  • The Model Context Protocol specification — modelcontextprotocol.io, the interface several of these agents now expose so a host application can call them without adopting their framework.
  • Demand figures on this page are this project's own paid measurements, recorded in its search_volume.json; repository behaviour was read from the projects' public documentation, not from any private source.

Method note

The slice run for this page recorded 8 of 8 sections pinned, 0 abstention(s) and 0 no-slice verdict(s) — and the page still carries no code excerpt, which is a recorded finding rather than an omission. Every one of the eight pins resolved to a single generic asset: a web-search container class whose name collides with ordinary search vocabulary, because the matcher's local rule accepts a symbol whose name is a substring of the keyword and every phrase in this lane contains "research". Eight sections sharing one search container is a substring collision, not section-level evidence, so the house rule for an unpinned section applies and the page is written from the public sources above. No code, batch fingerprints, auction data or internal hosts appear on the page, so there is nothing here that has to be asserted verbatim.