Firecrawl Alternative: When to Replace a Crawl Pipeline
A Firecrawl alternative is a replacement decision, not a feature list. Five criteria decide it: whether the renderer executes the JavaScript a page needs, whether extraction returns the main content or the page chrome, where throughput limits and cost land, what each run records, and what the data-handling rules allow. Answer all five on your own corpus before you compare any product.
Short answer: A Firecrawl alternative is a replacement decision, not a feature list. Five criteria decide it: whether the renderer executes the JavaScript a page needs, whether extraction returns the main content or the page chrome, where throughput limits and cost land, what each run records, and what the data-handling rules allow. Answer all five on your own corpus before you compare any product.
Key takeaways
- Replace a crawl-and-extract pipeline when one criterion fails outright — not when a rival's feature list looks longer. An engine that renders nothing on your pages, or an extractor that returns navigation instead of article text, fails whatever else it does well.
- The migration cost sits in the carry-over, not in the cutover: the URL set, the extraction selectors, the render policy, the rate-limit budget and the copies of already-fetched documents.
- Verify on twenty of your own URLs before you sign anything. A vendor benchmark measures the vendor's corpus; your pages, your render policy and your robots obligations are what the replacement has to survive.
firecrawl alternative: the question behind the query
Search for a replacement for a crawl-and-extract service and you are usually bundling four different decisions into one word. Coverage: a page technology your current engine cannot render — a client-side application, an infinite scroll, a consent interstitial. Quality: the extraction returns the page furniture as often as the article. Economics: the per-page price or the request ceiling stops fitting the volume you actually run. Governance: where fetched bytes are stored, for how long, and under which terms you were allowed to fetch them.
Those four reasons resolve into five criteria, and the same five apply whether you are moving to a hosted crawl API, a self-hosted renderer with your own extractor, or a browser fleet you operate. Write your answer to each before you open a product page.
- Rendering. Does the engine execute the JavaScript a page needs before it extracts, and can you choose per request whether it does — or is rendering always on and always billed.
- Extraction quality. Given the rendered document, does the extractor return the main content, or everything between the first and last byte. Measure it, do not read the marketing.
- Limits and cost model. Requests per minute, concurrent jobs, page size and the unit you are charged in — per page, per token, per compute second, or per seat.
- Record. What does one run leave behind that you can query later: the URL, the status code, the engine that served it, the extracted size, the error when it failed.
- Data handling and compliance. Where do fetched documents live, how long are they retained, can you pin a region, and does the service respect the fetch rules your policy encodes.
The rest of this page takes each criterion, then turns the set into a checklist for the move itself. If the context around the fetch — what you keep, what you compress, what a later agent is allowed to assume — is the real problem rather than the fetch, start from context engineering as a discipline instead; this page is only about replacing the layer that goes and gets the bytes.
agentic search: what the fetch pipeline actually does
A production crawl pipeline is two phases and a policy layer. Discovery turns a query or a seed list into candidate URLs; acquisition fetches each URL, renders it if needed and converts it to a text format; and a policy layer decides which backends are allowed, in what order, and what is recorded about each attempt. Most replacement conversations go wrong because they compare only the acquisition half, while the pipeline's real behaviour — which engine ran, what it fell back to, why a URL was skipped — lives in the discovery and policy halves.
The excerpt below is the search-backend dispatch from our own stack, cut from the slice body. It resolves one effective backend from configuration, tries the configured engine, and records what actually served the request rather than what was requested: the engine is a value, not a code path.
# backend/smartgate/modules/search/algorithm.py — source lines 119–167 (search backend dispatch)
async def search(self):
import time
self.start_time = time.time()
backend = self.settings.resolved_backend()
self.effective_backend = backend
try:
if backend == "firecrawl":
rows = await self._search_firecrawl()
self.result_container.extend("firecrawl", rows or [])
elif backend == "searxng":
rows: list[dict] = []
try:
rows = await self._search_searxng()
except Exception as e:
logger.warning("SearXNG failed: %s", e)
if self.settings.searxng_fallback:
self.result_container.add_unresponsive_engine(
"searxng", str(e)
)
else:
raise
if rows:
self.result_container.extend("searxng", rows)
self.effective_backend = "searxng"
elif self.settings.searxng_fallback:
logger.info(
"SearXNG returned no rows; falling back to DuckDuckGo Lite"
)
try:
fb_rows = await self._search_duckduckgo_lite()
self.result_container.extend("duckduckgo", fb_rows or [])
self.effective_backend = "duckduckgo"
except Exception as fb_e:
logger.warning("DuckDuckGo Lite fallback failed: %s", fb_e)
self.result_container.add_unresponsive_engine(
"duckduckgo", str(fb_e)
)
else:
logger.warning(
"SearXNG returned no rows (searxng_fallback=false)"
)
else:
rows = await self._search_duckduckgo_lite()
self.result_container.extend("duckduckgo", rows or [])
except Exception as e:
logger.warning("Search backend '%s' failed: %s", backend, e)
self.result_container.add_unresponsive_engine(backend, str(e))
return self.result_container
# backend/smartgate/modules/search/algorithm.py — source lines 169–221 (search backend dispatch)
async def _search_firecrawl(self) -> list[dict]:
"""Firecrawl Cloud / 自建 v2:POST {api_base}/search"""
import httpx
base = self.settings.firecrawl_api_base.rstrip("/")
url = base if base.endswith("/search") else f"{base}/search"
headers = {
"Authorization": f"Bearer {self.settings.firecrawl_api_key}",
"Content-Type": "application/json",
"User-Agent": UA,
}
body: dict[str, Any] = {
"query": self.search_query.query,
"limit": self.search_query.max_results,
"sources": ["web"],
"scrapeOptions": {
"formats": ["markdown"],
"onlyMainContent": True,
},
}
async with httpx.AsyncClient(timeout=60.0) as client:
resp = await client.post(url, json=body, headers=headers)
resp.raise_for_status()
payload = resp.json()
data = payload.get("data")
out: list[dict] = []
if isinstance(data, list):
for item in data:
out.append(self._normalize_hit(item, "firecrawl"))
elif isinstance(data, dict):
for item in data.get("web", []) or []:
out.append(self._normalize_hit(item, "firecrawl"))
for item in data.get("news", []) or []:
out.append(self._normalize_hit(item, "firecrawl"))
return out[: self.search_query.max_results]
def _normalize_hit(self, item: dict, engine: str) -> dict:
snippet = (
item.get("markdown")
or item.get("description")
or item.get("snippet")
or item.get("content")
or ""
)
return {
"title": (item.get("title") or "").strip(),
"url": (item.get("url") or "").strip(),
"content": snippet.strip()
if isinstance(snippet, str)
else str(snippet),
"engine": engine,
}
Three decisions in that excerpt are worth copying into whatever replaces your current stack, and
they are decisions rather than code. First, one resolved backend per run: effective_backend
starts empty and is set to whichever engine actually returned rows, so a run's record can say
"asked for SearXNG, served by DuckDuckGo Lite" instead of reporting the request and hiding the
fallback. Second, failure is per engine, not per run: the SearXNG failure is caught, bookkept
as an unresponsive engine and — only when the fallback switch allows it — replaced by the next
option; an exception in one backend is data, not an outage. Third, the last branch is a default,
not an accident: when no backend matches, the code still calls the fallback rather than raising,
which is what keeps a mis-set configuration from stopping a scheduled crawl.
For a migration, the engine belongs in configuration and the record should name the engine that served the call. A pipeline whose engine is hard-coded in the fetch loop passes every staging test and then fails silently the first time a page needs a renderer the code does not have. When that loop becomes an agent is a different question, taken up below; here the engine choice only has to survive the move.
url to markdown: the conversion contract you must not lose
Whatever fetches the page, the artefact your application reads is usually markdown or plain text. That conversion is the contract a replacement can break most quietly, because a worse converter still returns something. Five properties are worth checking on your own pages rather than in a changelog.
- Main-content detection. Does the output contain the article, or the header, the cookie banner, the recommendation rail and the footer as well. Wrap the output of twenty URLs and read where the article starts.
- Structure preservation. Headings, lists and tables should survive as markdown; a table flattened to a paragraph is content loss even though the text is present.
- Code and preformatted blocks. Fenced code must keep its fences and its characters; an extractor that strips indentation changes the meaning of code it returns.
- Link and image handling. Absolute URLs, kept or dropped images, and what happens to a relative link when the base is lost.
- Metadata. Whether the canonical URL, title and language come out with the text, or whether you have to recover them from the fetch record.
The pinned excerpt for this section is honest about what it is. Rule A matched this phrase in our own repository on the token down — the longest token of "url to markdown" is markdown, and the matcher's local candidate was a keyboard handler in the dashboard's command surface, where those letters actually occur. It is the consuming client, not a converter, and it is reproduced because the matcher pinned it and the page does not silently drop a pin it cannot use.
# components/dashboard/search-command.tsx — source lines 24–29 (dashboard search command handler)
const down = (e: KeyboardEvent) => {
if (e.key === "k" && (e.metaKey || e.ctrlKey)) {
e.preventDefault();
setOpen((open) => !open);
}
};
The handler is a small thing: a key chord that opens the dashboard command palette, with the platform's own modifier key. Its relevance to a migration is exactly its location. Conversion output does not end at a file — it lands in a client that has to search it, summarise it and show it. A replacement that changes the output shape (a different heading style, tables flattened, front matter dropped) can break that client without breaking a single fetch. Freeze a sample of current output, run the same twenty URLs through the candidate, and diff the two documents rather than the two summaries. The format itself belongs to the URL to markdown conversion contract; this page only asks whether a replacement preserves it.
agentic retrieval: when the pipeline becomes a loop
A single fetch is a pipeline; a fetch that reads its own result and decides what to fetch next is a loop, and the two have different failure modes. A loop needs a stop condition that is not "the model said it was done" — a maximum step count, a token budget, a wall clock — and it needs a ceiling that is enforced per run rather than reported afterwards. If the thing you are replacing already runs a loop, the replacement must be able to express one; if it cannot, you will rebuild the loop in your own code on top of a single-request API, which is a legitimate design but one you should choose knowingly.
Three questions separate the candidates here. Does one logical request carry state across steps, or is every call independent and the state yours to hold. Can a step be retried without re-fetching or re-billing the steps that already succeeded. And is there a per-run budget — pages, tokens or seconds — that the platform enforces before the call rather than after it. A candidate that answers "yes, yes, yes" is more expensive to run and much cheaper to operate; a candidate that answers "no" is a library, and the loop becomes your problem.
The loop is also where cost stops being linear. A loop that re-fetches a page because a downstream step failed twice pays three times for the same bytes, and the only defence is a record that says which fetch ids were already paid for. That record is the same object as the crawl log, which is why the fourth criterion — what a run leaves behind — is not an observability nicety but the thing that makes retries safe. What a retrieval loop is for, and how it differs from a one-shot search, is worked through on agentic retrieval.
context compression: the tokens the fetch creates
Every page you fetch becomes tokens somebody eventually pays for, and the two costs are decided in different places. The fetch decides how many bytes arrive; the pipeline decides how many of them reach a model. A replacement that returns larger, noisier documents — the whole page instead of the article, tables as prose, boilerplate repeated on every URL — moves cost downstream where it is harder to see, because the fetch invoice looks unchanged while the model invoice grows.
The measurable version of this is a compression ratio on a fixed corpus. Take twenty URLs, run them through the current pipeline and the candidate, and record three numbers per URL: the raw rendered size, the extracted size, and the size after whatever de-duplication you apply. A candidate that extracts cleanly usually wins the second number and the third, and the difference compounds across a loop that re-reads the same material on every step.
Migration carries two artefacts here: the extraction configuration — selectors, main-content heuristics, the list of elements to drop — which has to be reproduced or re-derived, and the compression policy, if the fetch layer also summarises or truncates, because a policy that silently drops the tail of a long document is a behaviour change, not an optimisation. What compression does to meaning is the subject of context compression; the migration point is narrower and blunt: copy the settings, then re-measure the ratio.
llm memory: what the fetched bytes are allowed to leave behind
A crawl pipeline accumulates a store of documents that is easy to forget about and expensive to defend. The questions a replacement has to answer are retention questions, and they are asked about the new pipeline before it runs, not after the first audit.
- What is stored, and where. Raw HTML, rendered DOM, extracted markdown, or only vectors — each is a different data asset with a different sensitivity.
- How long, and who can read it. A retention window enforced by deletion rather than by intention, and a list of the service accounts, engineers and vendors who can read the store.
- What the source permitted. Robots rules, terms of service and any contractual limit on the pages fetched; a replacement that ignores those moves the risk, it does not remove it.
- What is exported when the pipeline is retired. The retention clock does not travel with you, so the export happens before decommissioning, not after.
The last item is the one teams skip, and it is the one that turns a migration into a data incident: the old service's copies go away when the account is closed, taking with them documents your own retention policy promised to keep. What a language-model system is allowed to remember is on LLM memory. This page's version is a checklist item: inventory the store, export what your policy covers, then close the old account.
memory agent: when the crawler writes as well as reads
A crawler that only reads is a pipeline. A crawler whose results are written into a memory store that a later agent reads is a different system, with a different failure mode: a bad fetch does not just produce a bad document, it produces a fact that every subsequent run will retrieve. The replacement question here is whether the new pipeline can express the write path as a reviewed operation rather than an automatic one.
Three controls make that possible, and a replacement should be judged on whether it can support them rather than on whether it ships them. A write is attributable — the record says which run and which URL produced the memory, so a wrong fact can be traced and removed. A write can be held — a queue a human or a rule approves before the memory becomes retrievable. And a write can be reverted — deleting a memory by its source. A pipeline that writes straight into an index has none of these, and the first hallucinated fact is permanent.
The pattern itself is the subject of the memory agent pattern. For a migration, the criterion is narrower: does the candidate's record link a stored document to the run that fetched it. If it does, you can re-index after the move; if it does not, the safest plan is to re-fetch rather than to import a store you cannot attribute.
agent memory systems: the five criteria as a migration checklist
The five criteria become useful when they are turned into an inventory. Nine things move during a crawl-pipeline replacement, and each one can be made explicit before the cutover day rather than discovered during it.
| What to carry | Why it breaks a cutover | How to make it explicit |
|---|---|---|
| The URL set and the seed list | A dropped seed silently shrinks coverage | Export the seed list and diff the first full crawl against the old one |
| Extraction selectors and drop rules | A re-derived extractor changes what the text contains | Keep the configuration in version control, not in the old service's UI |
| Render policy per host | Treating a static page as an app multiplies cost, and the reverse loses content | Record which hosts needed rendering and re-check the flags |
| Rate limits and the politeness budget | A new pipeline with a higher default hammers sites you used to crawl gently | Re-assert the per-host limits from the old configuration |
| Credentials and API keys | Applications and schedulers authenticate with them | Issue the new keys in the same window and retire the old set |
| The crawl record and its retention | The new pipeline starts with no history, and responses still quote the old one | Export the run history your policy covers before anything is closed |
| Stored documents and vectors | A new index starts empty while consumers still expect answers | Decide re-fetch versus import, and re-index rather than guess |
| Extraction regression cases | They encode what "correct" meant on the old pipeline | Point the existing cases at the new output and diff the pass set |
| Runbooks and alerting | On-call looks at a surface that no longer exists | Rewrite the pages on-call actually opens |
Sequence the move as inventory, shadow, canary, cutover, retire. Shadow is the step that pays for itself: run the candidate alongside the current pipeline over the same URL set and diff the two records and the two extracted documents. That comparison finds the render-policy mistakes and the extraction regressions while both systems are live. And write the rollback trigger down before the window opens, not during it.
Applying the five criteria to this stack
SmartGate is the layer this page describes from the inside, so the honest presentation is against the same five criteria rather than against a competitor's feature list. The acquisition path exposes one endpoint, and the record and the enforcement point are the same object: a call writes an audit row, and the per-key window and the per-team budget read that row. That is what makes the fourth criterion cheap to answer here and expensive to answer on a pipeline whose only record is an aggregate dashboard.
The plan decides operational limits rather than features. Monthly token caps across the four tiers are 2M, 20M, 100M and 200M+; requests per minute per key are 120, 300, 600 and 1200; audit-log retention is 7, 30, 90 or 180 days; and a team can hold 2, 10, 30 or 9999 keys. Those are the numbers a replacement decision has to fit around, and the pricing page is the authoritative table — read the current values against your own retention obligation and your own peak traffic before treating anything above as settled.
On carry-over, be equally plain. Applications move by repointing a base URL and re-issuing a key, and the regression suite you already run is the acceptance test. If one of the five criteria fails for you — a renderer that cannot execute what your pages need, a retention window shorter than your obligation, an export your policy cannot read — the correct answer is to keep what you have, or to run both pipelines until the gap closes. A replacement that fails a criterion you actually have is a downgrade however good its documentation is.
Frequently Asked Questions
Limitations
This page is a decision framework, not a comparison. No product named or implied here is ranked, scored or benchmarked by us, and none of the criteria come from a load test we ran against a competitor: they are the questions a replacement decision has to answer, and the evidence that fills them in has to be yours. Treat the five criteria as an inventory to complete, not as a score out of five, and weight them by what your own pipeline actually fails at.
The page quotes two code excerpts and both are narrow. The 2:1 excerpt is a token-level match
rather than a conversion routine — the matcher pinned a client handler whose symbol contains the
letters down — and it is included because the method below records it, not because it
demonstrates URL-to-markdown conversion. The 8:1 excerpt is a backend dispatch from one
repository and does not describe any other product's internals.
The plan figures quoted above are operational limits rather than a feature list. Caps, per-key
rates, retention windows and key counts change with the plan, and a compliance decision should read
the current /pricing table rather than this page. Where a trade-off exists it is stated as one:
running two pipelines for a month costs double and buys the only comparison worth anything, your
own corpus.
Sources
- Firecrawl's own documentation — docs.firecrawl.dev — for the request shape a hosted crawl-and-extract service exposes, including search, scrape and the formats it returns.
- The Model Context Protocol specification — modelcontextprotocol.io/specification, for how tool traffic is described and returned when the fetch is one tool among several.
- robots.txt, RFC 9309 — rfc-editor.org/rfc/rfc9309, for the fetch rules a replacement pipeline has to keep honouring rather than relax.
- Google's documentation on JavaScript rendering and crawling — developers.google.com/search/docs/crawling-indexing/javascript, for why a page that needs a renderer is a different acquisition problem from a static one.
- The HTML Standard's section on scripting — html.spec.whatwg.org, for what a renderer is executing on a client-rendered page.
- The Common Crawl project — commoncrawl.org, as the reference example of a large crawl pipeline whose record and politeness policy are public.
- Scrapy's documentation — docs.scrapy.org, for the shape of a self-hosted acquisition pipeline with its own throttle, retry and item-pipeline configuration.
- Playwright's documentation — playwright.dev, for the browser automation a self-hosted renderer is normally built on.
- Demand figures on this page are our own measurement (DataForSEO Google Ads, United States, 12-month window, measured 2026-10-01), recorded in this project's keyword and brief files.
- Product behaviour and the plan table: read from the product source at the revision pinned in this
project's
pipeline_results.json, read-only, with the plan figures re-verified against the live pricing page on 2026-10-01.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
|---|---|---|---|---|---|---|
| 1 | agentic search | Search |
backend/smartgate/modules/search/algorithm.py |
119–167, 169–221 | rule A L2 → slot-proof | 5d88cd6aed7c |
| 2 | url to markdown | down |
components/dashboard/search-command.tsx |
24–29 | rule A L2 → slot-proof | 42939cf66801 |
Method note
This page carries two code excerpts and six sourced sections. The slice matcher pinned
2 of 8 sections for this page (0 abstention(s), 6
no-slice verdict(s)): rule A found a unique local symbol for agentic search and for
url to markdown, and no unique symbol for the remaining six section phrases, which were then
written from sources. The two excerpts were cut out of the slice body and re-asserted against it
byte-for-byte before publication, so nothing here is transcribed by hand.
Two readings belong in the record rather than in a footnote. The agentic search pin is a level-2
local match on a search-backend dispatch, and the excerpt is quoted because that is what the
matcher pinned — it is one repository's shape, not a specification. The url to markdown pin is a
token-level match: rule A's unique local candidate was the symbol down, a keyboard handler in a
dashboard client, and that is stated in the section above rather than hidden. Section keyword
phrases echoed in the headings come from this project's own paid measurement run, not from a
third-party tool. No batch fingerprints, auction data or internal hosts are transcribed, so there
is nothing here that has to be asserted verbatim.