SmartGate

URL to Markdown: What a Model-Ready Page Has to Keep

URL to markdown is the conversion contract between a web page and a language model, and it is a selection problem before it is a syntax problem. Fetch the document, decide which part of it is actually the page, drop the chrome, and serialize what remains so the structure still carries meaning.

Short answer: URL to markdown is the conversion contract between a web page and a language model, and it is a selection problem before it is a syntax problem. Fetch the document, decide which part of it is actually the page, drop the chrome, and serialize what remains so the structure still carries meaning. Four decisions settle whether the output is usable: main-content selection, boilerplate removal, preservation of tables, code fences and link targets, and a token budget measured before the text reaches the model.

Key takeaways

  • Selection comes first. Tag stripping keeps everything inside the tags, which is the wrong half.
  • Chrome is the default failure. Navigation, banners and footers repeat on every page and can outweigh the article they surround.
  • Structure is content. A flattened table or an unfenced code block is worse than no extraction.
  • Links have to survive as absolute targets, or the model cannot cite and the agent cannot follow.
  • Budget before you send. A converted page is a token commitment; measure it per task, not per page.
  • Do this next: take ten pages your agents already read, convert them, and count what the chrome was costing you in tokens.

url to markdown: what the conversion actually has to keep

Ask what the job is and the answer sounds mechanical: fetch the URL, read the HTML, turn the tags into markdown. The mechanical part really is small, and it is not where the work is. What decides whether the output is usable is four judgements about meaning, taken before any tag becomes syntax.

What is the page. A fetched document is mostly not content: navigation, cookie notices, newsletter prompts, related-article rails, comment widgets, a footer sitemap. On a typical documentation page the chrome can outweigh the article it surrounds, and every token of it is paid for twice — once when the model reads it, and again each time it lands in a retrieved chunk.

What structure has to survive. Markdown's value is that it carries structure cheaply. A heading becomes a heading, a list stays a list, a table keeps its columns, a code block keeps its fences and its indentation. Flatten any of those into a run of words and the relationships that made the page useful are gone: numbers without their column header, a snippet without its language, steps without their order.

Where the links point. Relative hrefs mean nothing outside the document that carried them, so a converter resolves every link against the base URL, keeps the anchor text, and decides explicitly what to do with the navigation links it otherwise drops. If citations matter downstream, targets have to be absolute.

What fits in the budget. Conversion sits upstream of a context window. A long page is a real token commitment, so the output should be measured — and, when necessary, split on section boundaries — before the request is made, not after it fails.

Conversions fail in two opposite directions, and both are silent. Under-extraction drops something structural — a table flattened to whitespace, a decision list collapsed into one sentence, a caption separated from its figure — and the model answers confidently from what is left. Over-extraction inflates the page with chrome until the actual answer sits somewhere in the middle of a document that is mostly menus. Neither raises an error, which is why an extractor needs a check of its own: convert a page you know well and read what came out.

That is the loop: fetch, select, serialize, budget. Which of those steps you own and which you buy is a separate question, and the place this step sits inside a wider practice, where context is assembled, reduced and remembered, is mapped on context engineering.

The excerpt below is what the slice matcher returned for this section, and it is more useful as a warning than as evidence. The pin is a name-level match: the symbol down occurs inside the keyword url to markdown as a substring of the word markdown, which is a coincidence of spelling rather than a statement about behaviour. The provenance table records the match level for exactly this reason. It is also a fair specimen of the problem this section describes — six lines of a keyboard-shortcut handler from a dashboard search dialog, a UI affordance that a converter keeping everything would cheerfully ship into a model's context. Nothing else on this page rests on it; the claims come from the sources listed at the end.

# components/dashboard/search-command.tsx — source lines 24–29 (down)
const down = (e: KeyboardEvent) => {
      if (e.key === "k" && (e.metaKey || e.ctrlKey)) {
        e.preventDefault();
        setOpen((open) => !open);
      }
    };

agentic search: results are urls, and one call can already return markdown

An agentic search call is the front door of this pipeline, and it returns two different things that are easy to confuse: snippets, and the URLs behind them. A snippet costs a few hundred tokens and often answers the question or rules the page out. A URL is an invitation to pay for a full conversion.

The discipline that follows is cheapest-first. Search, read the snippets, rank the candidates, then convert only the pages the loop actually depends on. An agent that converts every result at every step spends its budget on pages whose snippets already answered the question.

The conversion policy can be set at request time rather than after the fact. The excerpt below builds the request body for a fetching backend and asks it for one source type, a result limit, and — the part this page cares about — markdown with main content only. The caller therefore receives a converted page rather than a full document, and does not have to guess which of the two it is holding.

# backend/smartgate/modules/search/algorithm.py — source lines 180–188 (Search)
        body: dict[str, Any] = {
            "query": self.search_query.query,
            "limit": self.search_query.max_results,
            "sources": ["web"],
            "scrapeOptions": {
                "formats": ["markdown"],
                "onlyMainContent": True,
            },
        }
# backend/smartgate/modules/search/algorithm.py — source lines 194–204 (Search)
        data = payload.get("data")
        out: list[dict] = []
        if isinstance(data, list):
            for item in data:
                out.append(self._normalize_hit(item, "firecrawl"))
        elif isinstance(data, dict):
            for item in data.get("web", []) or []:
                out.append(self._normalize_hit(item, "firecrawl"))
            for item in data.get("news", []) or []:
                out.append(self._normalize_hit(item, "firecrawl"))
        return out[: self.search_query.max_results]

Two details in that excerpt matter downstream. The response is handled in both shapes the backend can return — a list of items, or an object carrying separate web and news collections — and the result is truncated to the requested limit before it leaves the method, so a backend that ignores the limit cannot flood the caller. The search container behind it dispatches across three engines, SearXNG's JSON API, a Firecrawl API and DuckDuckGo Lite, so one engine failing degrades the result set instead of failing the run.

The trade-off is worth stating plainly: the more the fetch layer does for you, the less control you have over the token shape of what comes back. Snippet-only results are cheap and shallow; converted pages are deep and expensive; a loop that cannot tell which one it is holding will misplan both.

firecrawl alternative: renting the fetch step or running it yourself

A generic scraping service sells the fetch step as a product: a URL in, a rendered page out, with browser execution, proxy rotation and per-site extraction rules maintained for you. That is often the right first move. The question this page adds is narrower — which half of the step are you buying, and what does it leave you owning?

The HTML-to-markdown conversion is the commodity half. Open extractors and serializers cover it, and the heuristics that decide what "main content" means are public and mature. The expensive halves are browser execution for pages that need script to render, and the operational shell around the fetch: retries, rate limits, robots handling, a cache and a budget.

Four questions separate the options faster than any feature list:

  1. Does the target set actually need a browser? Documentation, changelogs, blogs and most reference sites are static. Measuring that before buying browser execution is the cheapest saving available in this pipeline.
  2. What does per-page pricing do at your volume? A price that is trivial at a thousand pages a month is the dominant cost at a hundred thousand, and it grows with the corpus rather than with the value the corpus returns.
  3. Do you need markdown specifically, or would plain text do? If the consumers are embeddings only, structure is less valuable; if a model reads it and cites it, structure is the point.
  4. What happens on a refusal? A 403 or a rate limit that surfaces as a page-shaped 200 is worse than an error, because it enters the pipeline as content and is only discovered much later.

The largest saving available is usually not in either choice but in front of them. A conversion layer that caches by URL and content hash turns a repeated corpus into one fetch per change, which matters far more than the difference between two vendors' per-page prices when the same pages are read by several tasks. Pair that with a per-task token ceiling and the fetch bill becomes predictable even while the corpus grows, because the ceiling is set by the work rather than by the index.

What is genuinely durable here is not the converter. It is the cache keyed by URL and content hash, the per-task token budget, and the record of what was fetched and when — which is why Firecrawl alternative is a question about operating a fetch layer, not about picking a library.

agentic retrieval: quality starts at the extraction boundary

Retrieval can only rank what the extractor handed it, so the boundary between the two is where most quality is won or lost.

Chunk on structure, not on length. A fixed character window cuts a table in half and separates a step from the heading that explains it. Chunk at heading boundaries, keep the heading path inside each chunk, and overlap only the margins.

Boilerplate is a retrieval bug before it is a token bug. The same navigation, banner and footer appear on every page of a site. Chunked naively they become near-identical vectors scattered through the index, so a query about the site retrieves the chrome instead of the article. Main-content selection is therefore a ranking improvement, not only a saving.

Tables need their headers. A block of rows without the header row is a set of numbers with no units. Split a large table by row groups and repeat the header in each chunk.

Code needs its fences. A chunk that begins in the middle of a function looks like syntax and means nothing; if a function matters, keep it whole even when that makes the chunk uneven.

Overlap is a cost, not a safety net. The usual advice to repeat the last few sentences in the next chunk is worth paying for only when a boundary is genuinely lossy; repeating a tenth of the document in every chunk inflates the index and the context at once. Where a heading already restates the subject, the overlap buys nothing.

Identifiers belong on every chunk. The canonical URL, the section path and the fetch time are what a citation is made of, and they cost a few tokens each.

How those chunks are stored, scored and recalled once they exist is agentic retrieval; this page ends its responsibility at the text and the identifiers that leave the converter.

context compression: shrink the converted page, keep the answer

A converted page that is too large has two fates: it is dropped, or it is compressed. Compression is the better default, provided it happens before the model call rather than inside it.

The order that works is deterministic first. Deduplicate, because the same paragraph quoted on five fetched pages should be paid for once. Then filter by relevance to the actual question, lexically or by embedding score, keeping the passages that clear the bar. Then truncate, and truncate with an explicit marker so the model can distinguish a cut from an ending.

What must not be compressed away is short and specific: numerals, units, dates, table headers and code. A summary that turns a precise limit into "a generous allowance" has destroyed the only part of the sentence a reader needed, and it did so invisibly.

Measure in the tokenizer you will actually pay for. Character counts and word counts are useful proxies for a pipeline's own monitoring, but the budget is denominated in the units the model bills, and those differ between models by more than the margin people assume. Counting before the call, in the right units, is what makes a ceiling enforceable rather than aspirational.

Two cautions. A model-written summary costs a call, is non-deterministic, and can introduce a claim the page never made, so it belongs after the cheap passes rather than instead of them. And a per-task budget beats a per-page rule: the number that matters is the total a task may spend, which is what turns "our context is too big" from an opinion into a measurement. The mechanical reductions — dedup, relevance filtering, truncation, quota — are worked through on context compression.

llm memory: what to keep after the page has been converted

Once a page has been converted, it will be asked for again — by the same agent tomorrow, or by an unrelated task next week. The cheapest optimisation in this pipeline is not converting it twice.

A conversion cache needs three fields and one decision. The fields are a normalized URL, a content hash or the server's own validator such as an ETag, and the time of the fetch. The decision is how stale a stored copy may be before it is fetched again. A documentation page and a news article do not deserve the same answer, so make the freshness window a property of the source rather than a global constant.

Invalidation deserves the same explicitness as the fetch. A validator that comes back unchanged means the stored copy is still current whatever its age; a validator that has changed means re-convert, whether the copy is a day old or a year. Where the origin offers no validator at all, the honest fallback is a time-based rule plus a note that the copy may be stale, because reading yesterday's version of a live page is a silent error rather than a visible one.

Two practices keep that cache from becoming a liability. Store the markdown and its metadata, not the raw HTML: the conversion is deterministic from the source, so keeping both doubles the storage without adding information, and the canonical URL is what makes a re-fetch possible when the converter improves. And treat retention as a policy rather than a default, because converted pages can contain personal data and how long they are kept is a decision with a storage cost, a privacy dimension and an audit dimension. Where fetches run through a gateway, the plan's log retention — 7, 30, 90 or 180 days — is one input to that decision rather than the whole of it.

What a durable store of converted pages looks like once the questions stop being about fetching — what is retrieved by meaning rather than by key, and what may be written down at all — is the subject of LLM memory.

memory agent: the difference between a cache and a memory

A cache answers "have I seen this URL". Memory answers a different question: "what do I know because of it". The gap between those two sentences is where reading agents go wrong.

The unit of memory should be the fact, not the page. An agent that stores whole converted documents has built a search index, not a memory, and a later run will spend thousands of tokens to recover one sentence. Extract the claim, keep the URL and the fetch date beside it, and retrieval becomes cheap and checkable.

Provenance is what makes a stored fact falsifiable. A fact without a source cannot be re-checked when it looks wrong; a fact without a fetch date cannot be aged. Both are structural defects rather than housekeeping, and they have opposite symptoms: the first lets a stale fact survive indefinitely, the second makes it impossible to tell that it has.

Writes should be deliberate. An agent that writes memory without review will eventually persist its own mistake and then read it back as evidence, which is why the workable division is that automatic writes cover derived, reproducible data — the page, its hash, its section path — while durable claims about the world go through review.

How much of this a platform retains is part of a plan rather than a feature: audit retention runs 7, 30, 90 or 180 days by tier, which sets the outer bound of how far back a memory can be re-checked. The pattern itself is the subject of memory agent.

agent memory systems: the schema your extractor decides

Every memory system is downstream of an extractor, and the extractor's output is effectively a schema. Whatever it discards is discarded for every system built on top of it, later, by other people.

The fields worth keeping from a conversion are the boring ones: the canonical URL, the heading path of each block, the block type (paragraph, list, table, code), the fetch time, and a hash of the content. The first three make retrieval and citation possible; the last two make deduplication and staleness possible. Dropping the heading tree is the expensive mistake, because it is the only thing that tells a later chunker where a section began and what it belonged to.

Idempotency is the second decision. The same URL converted twice must not become two memories; a content hash is the natural key, and it makes a re-fetch cheap to detect — same hash, no write, only a newer timestamp.

The third decision is what counts as the source of truth. Storing the markdown is reproducible: the same input and the same converter produce the same text, so a bug in the converter can be found by re-running it over the same corpus. Storing only a model's summary is not, because the summary cannot be audited against anything except itself. Summaries are derived artefacts and belong beside the source, never instead of it.

A fourth field is worth the few bytes it costs: the extractor that produced the text, as a version string. When the selection rules change — a new container is recognised, a class of chrome is dropped — stored conversions from the old version are a different kind of object from new ones, and a version stamp is what lets a re-conversion sweep be scoped to the pages that need it instead of the whole corpus.

Once those three are settled, the higher-level questions — how recall is scored, how long each class of memory lives, who is allowed to write — become ordinary engineering instead of guesswork.

Where SmartGate fits

SmartGate is an MCP-native algorithm gateway, and one of its seven primitives is the step this page describes: smart_fetch fetches a public URL and converts it to Markdown. The other primitives cover aggregation search, context compression, duplicate removal, budget enforcement and team memory, and the pipeline templates built on them (smart_pipe) are what arrange research, read and remember as sequences rather than as separate products.

What the gateway adds around a conversion is the operational half. Every call resolves to a key and a team, is written as an audit row, and counts against the limits that key carries. Four tiers set those limits: monthly token ceilings of 2M, 20M, 100M and 200M+, per-key MCP request rates of 120, 300, 600 and 1200 per minute, audit-log retention of 7, 30, 90 or 180 days, and 2, 10, 30 or 9999 keys per team. Those are operational numbers rather than features, and the pricing page is the authoritative table — compare the retention window against the longest re-check window your own corpus needs before committing to a tier.

Frequently Asked Questions

Limitations

This page is about the extraction step, not about whether a page may be fetched. Robots rules, terms of service, licensing and copyright are separate questions with separate answers, and nothing here should be read as settling them.

Main-content extraction is heuristic and per site. A redesign that moves an article into a new container, or a page whose text is split across tabs, can change what an extractor selects without any error being raised — which is why the useful monitoring signal is a change in the token count of a converted URL, not a change in its status code.

No converter is benchmarked or ranked here, and no claim is made about the quality of any particular library or service. The two excerpts above are one name-level pin and one request/response window from the repository the matcher scanned; neither one is evidence about how well any converter performs.

The plan figures quoted above are operational limits rather than a feature comparison, they determine ceilings, retention and key counts rather than capabilities, and the pricing page is the authoritative table — a retention commitment should be read from there rather than from this page.

Sources

  • John Gruber — Markdown, the original syntax definition, for what structure a markdown document is expected to preserve.
  • CommonMark — spec.commonmark.org, for the current normative definition of the format the consumer renders, including tables and fenced code.
  • Mozilla — Readability, the extractor behind Firefox's reader mode, as the reference shape of main-content selection.
  • Trafilatura — github.com/adbar/trafilatura, for the extraction-then-serialization pipeline and its evaluation work on boilerplate removal.
  • Turndown — github.com/mixmark-io/turndown and html2text — github.com/Alir3z4/html2text, as the two widely used HTML-to-markdown serializers.
  • Firecrawl — docs.firecrawl.dev, for the options a hosted fetch backend exposes on a request, including the markdown format and a main-content-only switch.
  • Model Context Protocol — modelcontextprotocol.io, for how a fetched document is handed to a model through a tool call rather than pasted into a prompt.
  • Demand figures in this page are our own measurements: DataForSEO Google Ads, United States, 12-month window, measured 2026-10-01, recorded in this project's keywords.txt and search_volume.json.
  • Product behaviour and the plan figures were read from the product source and the published plan table, read-only; the pricing page is the authoritative table for every limit quoted above.

Method note

Two of the eight sections on this page carry a code excerpt, and the other six carry none. That is a finding rather than an omission: rule A needs a symbol whose name the section keyword identifies uniquely, and on the repository the matcher scanned that succeeded twice while returning no unique symbol for retrieval, compression, memory and the other six phrases, whose vocabulary also names generic helpers and types across a codebase. Sections without a pin are written from the published sources listed above, with no claim about how any implementation works.

Both excerpts were cut out of the slice body returned by the slice API and re-asserted byte-for-byte as substrings of that body before publication; the first line inside each fence records the file and the exact source lines, and the provenance table below records how each pin was matched — including the name-level match on the first section, which is why this page does not treat that excerpt as evidence. Product behaviour was read from the product source at the pinned revision, read-only. No batch identifiers, auction data or internal hosts appear anywhere on this page.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)
1 url to markdown down components/dashboard/search-command.tsx 24–29 rule A L2 → slot-proof 42939cf66801
2 agentic search Search backend/smartgate/modules/search/algorithm.py 180–188, 194–204 rule A L2 → slot-proof 5d88cd6aed7c

Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before publication. 2 of 8 sections pinned, 0 abstentions, 6 misses.