Deep Research API: Wiring Research Into Your Own Product
A deep research API is a long-running asynchronous job, not a chat call: you submit a question, get back an identifier, and collect a cited report minutes later by polling, streaming or a webhook. Integrating one well means deciding the job lifecycle, the quota, the failure path and the retention window before you decide the prompt.
Short answer: A deep research API is a long-running asynchronous job, not a chat call: you submit a question, get back an identifier, and collect a cited report minutes later by polling, streaming or a webhook. Integrating one well means deciding the job lifecycle, the quota, the failure path and the retention window before you decide the prompt.
Key takeaways
- It is asynchronous by construction. A research run takes minutes, so any endpoint that blocks on the answer will time out on exactly the questions worth asking.
- The job is the interface. Submit, status, result, cancel and an idempotency key are the whole contract; the report is just its terminal state.
- Quotas and rate limits are per key, per minute, so the caller needs a queue and a backoff, not a retry loop.
- Cache at the source, retry at the step. A URL-and-hash cache removes most repeat cost; a bounded retry policy handles the rest.
- Degrade, do not fail. A partial report with the sources that responded beats a timeout.
- Do this next: wrap the vendor call in your own job record — an id, a status, a cost counter and a stored result — before you wire a single prompt.
The deep research api contract: integrate it without freezing your product
The phrase covers a menu of things — a vendor endpoint, a self-hosted agent, an internal service over your own corpus — so the first useful move is to name the contract all of them must satisfy. A deep research api is a job interface. You submit a question with constraints, you receive an identifier, the work runs for minutes, and you collect a report at the end. If you model it as a request that returns an answer, you will build a caller that times out on precisely the questions that justify the call.
Five verbs are enough, and a service that omits any of them pushes the gap into your application code. Submit takes the question plus the knobs that matter (sources allowed, a cost ceiling, a deadline) and returns an id. Status reports one of a small terminal-or-not set — queued, running, waiting on input, done, failed, cancelled — and, crucially, how far along the run is. Result returns the report and its citation list once the run is terminal. Cancel stops a run whose answer is no longer wanted, which matters when the caller is a user who changed their mind. List lets an operator see what is in flight per tenant, which is the difference between a bill you can explain and one you cannot.
Two properties turn that contract from workable into operable. The first is an idempotency key on submit: a client that retries a timed-out submit must not start a second billed run, so the server returns the original job for a repeated key. The second is that the job carries its own identity — tenant, requester, budget and deadline — rather than inheriting them from an ambient session, because a research run outlives the request that started it by an order of magnitude.
That seam is where this page lives. Which research capability is worth wiring at all — the selection questions, the build-versus-buy call, the vendor comparison — is the subject of the AI research tool landscape. This page starts after that decision is made and asks only how the integration should be shaped.
openai deep research api: the asynchronous job shape you inherit
The clearest worked example of the shape is the one a large vendor already ships, because it settles several design questions for you. Deep research on the OpenAI platform is not a chat completion; it is a call to the Responses API with one of the deep-research models, and the documentation is explicit that the request should run in background mode and that a webhook can notify your service when the run completes. In other words, the vendor has already decided that the endpoint is a job, and your integration inherits that decision whether or not it fits your product.
Three details of that shape are worth copying deliberately. First, the output is a trace, not just an answer: the response exposes the search calls, the code-execution steps and the tool calls that led to the report, which is what lets a downstream UI show progress and lets a reviewer check the work. Second, polling is time-boxed — the platform keeps background response data only briefly so that polling works reliably, and it warns that this temporary retention is incompatible with a strict zero-retention posture. That is not a footnote; it is a constraint that decides whether a regulated tenant can use the feature at all. Third, the cost and latency dial is a tool-call ceiling: a parameter caps how many tool calls the model may make, and the documentation names it as the primary lever on cost and latency.
For the integrator, the lesson is to mirror the vendor's job in your own job record rather than exposing the vendor's ids to your callers. Keep a stable local id, store the vendor id beside it, translate their state machine into yours, and record the tool-call ceiling you applied. When the vendor changes a model or a field, only the adapter moves. How a research run is directed and reviewed from the human side — the questions a researcher actually asks — is worked through where this cluster covers the AI researcher's workflow; this page keeps to the plumbing.
gemini deep research api: plans, steps, and reconnecting to a run
A second vendor answers the same design questions differently, and the differences are instructive. Gemini's Deep Research is an agent reached through the Interactions API, and the documentation is blunt that it is asynchronous: the call must set background execution, returns an interaction id immediately, and is polled or streamed until it reaches a terminal state. A plain synchronous generation call cannot reach it at all. That is a useful reminder that "deep research API" describes a class of job interface, not one vendor's route.
The part worth borrowing is how progress is exposed. The run emits an ordered list of steps, so a caller can render "searching", "reading", "synthesising" rather than a spinner. Streaming is an option alongside polling, and — the detail that separates a demo from a product — the stream can drop and resume: the client keeps the last event id and reattaches from that point, because a long research run will outlive an idle HTTP connection. If your product shows progress, plan for the reconnect on day one; retrofitting it after the first dropped stream is expensive.
A second borrowable idea is plan review before spend. The agent can be configured to return a proposed research plan and wait for approval instead of running immediately, with the conversation continued across turns until the plan is accepted. For a product that bills research per run, a plan-approval step is the cheapest possible cost control: the expensive part happens only after a human has seen the direction. The counterweights are real — the plan turn is another round trip, and the run still has a hard wall-clock ceiling after approval. Where those trade-offs land for a report whose output is a document rather than an answer is the subject of reading an AI research paper; the integration question here is simply whether the plan turn is exposed to your caller or kept inside your own service.
deep research: what the endpoint is really doing for those minutes
The reason the interface is asynchronous is the work underneath it, and knowing the loop lets you design the status values a caller actually needs. A research run is an agentic loop: formulate queries, search, read the pages, keep the passages that matter, and repeat until a stop condition — a step or tool-call budget, a deadline — is reached, then synthesise the kept material into a cited report. Each cycle can open new pages, which is why a run that starts in seconds can run for many minutes and why its shape is a fan-out followed by a synthesis step.
That loop dictates what a good status surface reports. A single "running" flag is not enough, because the caller cannot tell a productive run from a stuck one. Report the stage (searching, reading, synthesising), the count of sources retrieved so far, and the budget consumed so far against the ceiling. Those three numbers let a UI show honest progress, let an operator notice a runaway run before it exhausts a monthly quota, and let a client decide to cancel a run whose early sources are already answering the question. A run whose only signal is a spinner will be cancelled at the wrong moment or left to burn.
The loop also explains the failure modes you will actually see. Searches return nothing useful; individual pages refuse to load or rate-limit the fetcher; the synthesised report is long but shallow because the loop stopped early. None of these is a clean error, and all of them should reach the caller as a degraded but usable result rather than as a failure, provided the report says what it could not reach. How the reading and synthesis half is organised once the sources exist is the territory of an AI literature review; the API-side point is that the status surface and the failure surface are the same design problem, and the job's own record is where both live.
deep research ai as an api: sync, async, or streaming — pick the shape the caller can hold
"Deep research AI" reaches a product in one of three interface shapes, and choosing wrongly is the most common integration mistake. A synchronous call returns the report on the same connection. It is only honest when the run is short and the corpus small — otherwise the caller's own timeout, load balancer idle limit or mobile network will sever the connection long before the report exists. A polling job returns an id immediately and the caller asks for status on an interval. It is the most portable shape and the one to default to, because it survives deploys, restarts and flaky networks, at the cost of a little latency and wasted polls. A streamed job pushes updates over a long-lived connection. It gives the best UX and is worth it when a human is watching, but it forces you to solve reconnection and event replay, so it is a layer on top of the job — never a replacement for it.
The deciding question is not the technology; it is what the caller can hold. If the caller is a background worker with a persistent store, a polling job fits. If it is a browser tab a person is watching, a stream is worth the complexity — as long as the job still exists server-side and can be re-read after the tab closes. If it is a webhook target, the job ends by notifying a URL your service owns, which is the cheapest way to decouple a long run from the request that started it, provided the notification is authenticated and re-drivable. The failure to avoid is a shape that only works while the initiating request stays open.
Two cross-cutting rules apply to all three. Make the job the source of truth and every transport a view over it, so a dropped stream or a closed tab loses no work. And give every shape the same deadline and budget, so switching transports cannot change the cost of a run. When the same question is asked repeatedly, the cheapest answer is often a cache rather than a fresh run — the distinction between a one-off report and a reusable knowledge asset is where a deep research AI stops being a feature and becomes infrastructure, and it is the line walked on deep research AI systems.
ai research tool behind your own api: quota, limits, and retries
Once research sits behind your product, its limits become your limits, and the caller has to be built for them. The binding constraint is usually a per-key, per-minute request allowance at the gateway, plus a monthly token or spend ceiling at the plan. A caller that ignores both fails the way overloaded systems always fail: a burst of retries against a limit produces more failures, which produce more retries, until the limit is exhausted and the product looks broken for every tenant at once. The fix is not a bigger limit; it is a queued, rate-aware caller.
The pieces are standard and worth naming so they are not skipped. Put submissions on a queue with a per-tenant fair-share, so one customer's bulk job cannot starve another's interactive one. Use bounded exponential backoff with jitter on retryable failures, and treat a rate-limit response as a schedule signal — it tells you when to come back — rather than as an error to hammer. Make retries safe by construction with the idempotency key from the submit verb, so a retried submit returns the original job instead of paying twice. And cap concurrency per tenant as well as rate, because ten concurrent long runs can exhaust an allowance that a rate limit alone would have spread out.
Design the ceiling to be read, not estimated. A per-run budget, a per-tenant daily cap and a per-key rate limit together cover the three ways a research feature runs away: one expensive run, one busy customer, and one noisy key. When the research capability is one of several behind a shared gateway, those limits belong in one enforcement point rather than re-implemented per caller — the same argument made for a deep research agent when it is wired into a loop rather than exposed directly.
ai web research: cache keys, retries, and graceful degradation
A research run is mostly web reading, so the cheapest reliability work happens at the fetch and cache layer, not in the model. Cache the fetch, keyed by source. The key should be a normalised URL plus a validator — a content hash, or the origin's own ETag or Last-Modified — and the stored value should be the cleaned text with its fetch time. A cache keyed that way turns a repeated corpus into one fetch per change and, just as important, makes a re-run cheap to re-check: unchanged validator means the stored copy is current regardless of its age. A second, smaller win is a negative cache for sources known to refuse, so the loop does not pay the same timeout on every run.
Retries belong at the step, not around the whole job. A single page that timed out should be retried a bounded number of times with backoff; a whole run that failed halfway should not be restarted from the top, because that re-pays for every page it had already read. This is only possible if the job's intermediate state is durable, which is another reason the job record — not the request — is the unit of work. Distinguish retryable conditions (timeouts, rate limits, transient network errors) from terminal ones (a refusal, an empty result set), so backoff is spent where it can help.
Degradation is the payoff. When some sources are unreachable or the budget runs out mid-run, the useful outcome is a partial report that names its gaps, not a failure: the caller gets the answer the material supported, the report lists what it could not reach, and the job is marked degraded rather than failed. A circuit breaker around an unreliable backend stops a struggling source from consuming the whole allowance, and a fallback path — a secondary reader, or a cached copy with a staleness note — keeps the run alive. The trade-off is honesty: a degraded report is only better than a failure if it says so up front, because a confident answer built on half the sources is worse than a clear "could not complete".
research automation: cost control, compliance, and data retention
The last integration concern is the one that decides whether the feature can ship at all, because it is about money and rules rather than code. Cost control starts at the run: a tool-call ceiling, a per-run budget and a stop condition that is not "the model said it was done". Across runs, a per-tenant monthly ceiling turns research automation from an unbounded commitment into a line item, and a per-run cost figure written into the job record is what makes the ceiling enforceable instead of aspirational. The single most effective saving is usually in front of the model: cached fetches and a cache reply for a repeated question remove whole runs before they start.
Compliance and retention follow from what a research run touches, which is other people's web pages and, often, your own corpus. Three decisions have to be made explicitly. What is the retention window for reports, intermediate fetches and the job's audit row — and does the shortest window the vendor offers satisfy it, or is a longer one required. Where does data reside, and does a run that ships fetched pages to a third-party model cross a boundary the tenant's policy forbids. And can the record be exported to an auditor, in a form that shows what was fetched, when, under which key and with what result — because an audit trail you cannot export is not one. These are the same questions an enterprise asks of any long-running data path, and they belong in the design review, not the incident review.
Two honest cautions belong here as well. A vendor's asynchronous mode can retain data for a short window specifically so that polling works, which may conflict with a strict no-retention requirement — so the retention posture has to be checked against the actual mode used, not the marketing page. And a per-key rate limit or monthly cap is an operational number, not a feature; it determines what the product can sustain, which is why the plan's published limits should be compared against the workload before a capacity promise is made.
What a programmatic research call gets back, read from our own REST surface
When research is a scheduled job rather than a conversation, the contract matters more than the
feature list. Ours is one endpoint per tool, read on 2026-10-07 from
backend/smartgate/api/v1/fetch.py and modules/fetch/models.py.
- One POST per tool, with a small explicit body. The fetch endpoint takes a URL and a timeout bounded to between five and 120 seconds, defaulting to 30. Bounded timeouts are the difference between a job that fails and a job that hangs until your scheduler kills it.
- Refusals are part of the contract, with a retry hint. Each call is checked against the same
per-team, per-tool rate limit the agent tools use, and exceeding it returns a 429 carrying a
Retry-Aftervalue. A caller that treats every non-2xx as fatal will retry a job that only needed to wait. - Every call is audited with its parameters. For a research API this is the difference between "the report is wrong" and "here is what was fetched, when, by which credential" — and it is the reason to prefer a gateway over embedding a scraper in each job.
- The payload carries quality, not just text: the extracted markdown, the page metadata, a quality score and a content length. A pipeline that ignores the score will happily summarise a cookie banner; a pipeline that checks it can retry, fall back, or drop the source and say so.
- The honest limit: research quality is bounded by extraction quality, and no amount of downstream synthesis repairs a fetch that returned navigation chrome. Score first, summarise second.
Where SmartGate fits
SmartGate is the intelligence layer between AI and the world — the place where the research primitives and the enforcement point are the same object. Its five capabilities cover exactly the surface this page describes: Research (search and fetch), Context (de-duplication and compression), Memory (team-level storage), Control (a hard budget ceiling) and Pipeline (research, read and remember arranged as sequences). Because it is MCP-native, a research run reaches it as a tool call, and every call is written as an audit row and counted against the limits that key carries — caller, route, tool, token count, latency and outcome on one record, which is what makes a long run observable and attributable rather than a black box.
The plan table sets the operational ceiling rather than the feature list: monthly token caps of 2M, 20M, 100M and 200M+, per-key request rates of 120, 300, 600 and 1200 per minute, audit-log retention of 7, 30, 90 or 180 days, and 2, 10, 30 or unlimited keys per team. Those are the numbers a capacity plan has to be checked against, and the pricing page is the authoritative table — compare the retention window against the longest re-check window your own workload needs before committing a tier, and treat the published limits as operational rather than as marketing.
Frequently Asked Questions
Is a deep research API just a slower chat completion?
No, and the difference is structural. A chat completion returns one answer on the connection that asked; a research API runs as a job that searches and reads for minutes and returns a report with its sources. The interface is asynchronous by necessity, which is why submit, status, result and cancel matter more than the prompt does. Treating it as a slow chat call is exactly how a caller ends up with a timeout on the questions that needed the research.
Should we poll or use a webhook?
Polling is the portable default: it survives restarts, deploys and flaky networks, and it keeps the job as the source of truth. A webhook decouples a long run from the request that started it and is the cheapest option when your service already owns an endpoint, provided the callback is authenticated and can be replayed. A stream gives the best experience when a person is watching, but it must sit on top of the job, never replace it, because a long research run will outlive an idle connection.
How do we stop research runs from blowing the budget?
Put a ceiling on the run and on the tenant. On the run: a tool-call or step cap and a wall-clock deadline, with the spent amount recorded on the job. On the tenant: a monthly ceiling and a per-tenant concurrency limit, so one customer cannot exhaust a shared allowance. Then remove work before it starts with a cache reply for repeated questions and a fetch cache keyed by URL and content hash. Where a vendor exposes a plan-approval step, use it, because the cheapest run is one a human declined.
What happens when some sources fail mid-run?
Return a partial report that names the gaps, and mark the job degraded rather than failed. Retry only the step that failed, with bounded backoff, and only for retryable conditions such as timeouts or rate limits, because restarting the whole run re-pays for every page already read. A circuit breaker around an unreliable source keeps one bad backend from consuming the allowance, and a cached fallback keeps the run alive. The one rule is that a degraded report must say so up front.
Limitations
This page is about the integration surface, not about the quality of any particular research engine, and it ranks and benchmarks nothing. It does not settle which vendor to use, and it makes no claim about the accuracy of any model's research output; those are separate questions with separate evidence. The failure and retention behaviour described here is a consequence of designing around long-running, source-reading jobs, and a given product may reasonably trade some of it away — for example, choosing a synchronous shape for a short, narrow search where the latency is acceptable. Nothing here is legal advice: retention windows, data residency and export obligations depend on the tenant, the jurisdiction and the data, and should be checked against the actual contract and the deployed configuration rather than against this page.
Sources
-
OpenAI — Deep research, for the Responses API model choice, background mode, webhook completion, the output trace of search and tool calls, the tool-call ceiling, and the polling-retention versus zero-retention caveat.
-
OpenAI Cookbook — Introduction to deep research in the OpenAI API, for the asynchronous job shape and the background-mode recommendation.
-
Google — Gemini Deep Research agent, for the Interactions API, required background execution, plan review before spend, streaming, and the in-progress to completed or failed state transitions.
-
Google — Background execution, for the interaction state machine and stream reconnection with a last-event id.
-
Google — Interactions API, for the interaction resource, observable execution steps and server-side conversation state.
-
Model Context Protocol — modelcontextprotocol.io, for how a research capability is reached as a tool call rather than pasted into a prompt.
-
Demand figures in this page are our own measurements: DataForSEO Google Ads, United States, 12-month window, measured 2026-10-02, recorded in this project's keywords.txt and search_volume.json.
-
Product behaviour and the plan figures were read from the product source and the published plan table, read-only; the pricing page is the authoritative table for every limit quoted above.
-
The section "What a programmatic research call gets back, read from our own REST surface" is our own implementation, read on 2026-10-07 from
backend/smartgate/api/v1/fetch.pyandmodules/fetch/models.py(origin/main). It states only what those files state.
Method note
This page renders no code excerpt, and that is a recorded finding rather than an omission. The slice
matcher pinned 8 of 8 sections for this page (0 abstentions,
0 misses), and all eight of those recorded pins resolved to the same symbol, Search, by a
level-2 name-level match: every section but one keyword contains the substring search inside the word
re-search, so rule A kept returning one generic search container instead of one symbol per section.
A name-level pin is a spelling coincidence and not a behavioural claim about deep research, so
quoting it would give the page the shape of a verified article with none of the substance.
Every statement above is therefore grounded in the published sources listed before this note and in SmartGate's own published plan limits, read from the product source and the pricing page read-only. The section vocabulary and its measured monthly volumes come from this project's own paid keyword run, recorded in this project's keywords.txt and search_volume.json. No batch identifiers, auction data, internal hosts or unverified-scan caveats appear anywhere on this page.