Agent Pipeline: What Runs, In What Order, and What Breaks
An agent pipeline is a fixed chain of steps where each step receives a named input, does one thing and returns one artefact the next step can consume. The order is written down; only the work inside a step is open-ended. That is what makes the chain testable, priceable and explainable — and it is the cheapest lane in this vocabulary to enter, at 880 searches a month and a difficulty of 0.
Short answer: An agent pipeline is a fixed chain of steps where each step receives a named input, does one thing and returns one artefact the next step can consume. The order is written down; only the work inside a step is open-ended. That is what makes the chain testable, priceable and explainable — and it is the cheapest lane in this vocabulary to enter, at 880 searches a month and a difficulty of 0.
Key takeaways
- A pipeline is a decision you make once, not a decision the model makes each turn. Move the order out of the model and into a list and three problems disappear at once: unbounded cost, un-auditable branches and a test suite that cannot assert anything.
- Steps should be boring, and short on purpose. Fetch, extract, deduplicate, compress, write. The open-ended part belongs at one or two seams, not inside every step.
- The step contract is four fields: what it received, what it returned, how long it took, and which branch it took. A step missing any of the four is a footnote in someone's incident review.
- Per-step timing is the cheapest observability there is. A total runtime tells you the run was slow; the offset of each step tells you which step owns the slowness, and whether the wait was in the step or in front of it.
- Retries are safe only where the step is idempotent. Reading and compressing are naturally repeatable; writing is not, so the write step is the one that needs a key rather than a position.
- Measure the lane before writing anything. The phrases here are cheap to win
(
agent pipelineat difficulty 0), which is unusual and worth using rather than assuming.
What an agent pipeline is, and why the shape matters more than the tools
A pipeline is the least glamorous object in an agent system and the one that decides whether the system can be operated. Its defining property is not intelligence; it is that the sequence is data. A list of steps, each naming a tool and its parameters, is a thing you can read, diff, version and explain to somebody who did not build it.
The alternative — one model call that decides everything — has exactly one advantage, flexibility, and four costs that arrive together. Cost becomes a function of the model's patience rather than of the work. Testing degrades into sampling, because the same input may take a different path. A failure has no location, only a transcript. And the run cannot be decomposed for accounting, so nobody can say which part of the bill belonged to which part of the job.
A pipeline buys all four back at the price of flexibility at the seams. The practical arrangement that survives production is not "pipeline or agent" — it is a pipeline whose two or three decision points are explicit. Everything mechanical is a step; the judgement is a step too, just a non-deterministic one, and it is named so that its cost and its failure rate are visible beside the deterministic steps.
Naming matters more than it sounds. If a pipeline's steps are called step_1, step_2, step_3,
the timing record is unreadable within a week. If they are called search, fetch, dedup,
compress, the record reads like a sentence, and a run that took twice as long has an obvious
suspect.
ai workflow automation: the boundary between a step and a decision
Automation projects fail at the same place: a decision was written as a step. The symptom is recognisable in three forms, and all three look like quality problems rather than design problems.
Non-deterministic retries. A step that "sometimes needs another go" is usually a judgement wearing a step's clothes — the retry is the model deciding it did not like its own answer. If a step needs a retry policy that says "up to three attempts", ask what the three attempts are choosing between. If the answer is "whether the output is good enough", the retry is a decision and belongs in an explicit step with a stated rule.
Branches nobody can audit. The moment a step's output determines which step runs next without that mapping being written down, the run's path stops being reconstructable. A branch is fine — a branch that exists only inside a prompt is not, because it cannot be listed, counted or aged.
A test suite that cannot assert. A step you can test has an input, an output and a rule for matching them. A step whose correctness depends on a model's opinion can only be sampled, which is fine as an acceptance test and useless as a regression test.
The working boundary, stated as a rule: if you can write the acceptance check, it is a step; if you cannot, it is a decision, and decisions get their own step, their own cost line and their own error budget. That single split is what lets a pipeline be measured at all — and it is why an agentic workflow keeps its decisions at the seams rather than inside every stage.
function calling: what a step actually receives and returns
A step's shape is a function signature, and the interesting part is what it is allowed to see. Three conventions keep a chain honest, and each one prevents a specific silent failure.
Input by name, from a named previous step. A step that reads "everything so far" cannot be
re-run alone, and it hides two bugs: a step ignoring its input, and a step recomputing what an
earlier step already paid for. Naming the producer (from_step) turns both into a visible edge.
One artefact out, in the shape the consumer expects. A step returning a partially-populated dictionary pushes its assumptions into every downstream step. The shape should be as narrow as the next step needs and no narrower — the compression step wants text, not the search result that found the text.
A status, not a boolean. success: false cannot distinguish "no data" from "the source broke"
from "this branch is not applicable today". A status with a reason is what lets the next step decide
whether to continue, degrade or stop — and it is the field that turns a failed run into a readable
one.
Two of those fields are easy to skip and expensive to add later, because adding them changes every record you already stored. The time to insist on them is the first step you write, not the thirtieth.
pipeline orchestration: offsets, idempotency and the second attempt
Orchestration is what happens around the steps, and in practice it is three decisions.
Where the time went. Record the offset at which each step started and how long it took, not just the run's total. The total is an alarm; the offsets are a diagnosis. A chain that suddenly takes twice as long with the same steps is usually one step waiting on a dependency, and that fact is invisible in every aggregate you can think of.
Which steps may be re-run. Reading a URL, extracting text, deduplicating and compressing are naturally idempotent — the same input produces the same artefact, and a second attempt costs only the work. Writing is the opposite: "add this to memory" applied twice stores two points unless the write carries an explicit key. So the retry policy is per step, not per pipeline, and the steps that write are the ones that need a key derived from the work rather than from the position in the sequence.
What a degraded step reports. The most useful status in a pipeline is "I ran, but not the way you asked". Compression that gives up and passes the original text through is a reasonable behaviour; reporting it with the same shape as a successful compression is not, because every consumer of that artefact then treats a fallback as a result. A fallback flag and an unchanged ratio turn a silent lie into a readable line.
Get those three right and the pipeline survives a bad day without a human reading every log line. Get them wrong and the pipeline is still running — it just cannot tell you which part of it is wrong, which is the state most teams discover only during an incident.
llmops: operating the chain after it ships
Running a pipeline is a different job from building it, and the difference is entirely about what is recorded. Four readings carry almost all of the value.
One row per call, with the caller attached. The row needs the identity that pays (key and team), the arguments that mattered, the outcome and a correlation id that groups the calls belonging to one task — the shape LLM observability describes field by field, and the surface the audit and compliance pillar reads. Without the correlation id, a task assembled from five calls across three services is not reconstructable; with it, the task is a query. Labelling the route (which surface the call arrived on) at write time rather than deriving it later from a route table is what keeps the label true after the first rename.
Spend attributable per key and per team. Aggregate spend answers "was this month expensive" and nothing else. The first runaway loop you meet will be one key, and the aggregate will not name it — which is the job token control exists to do.
Three counters instead of a score. How often the stopping rule fired late, how often a fallback branch was taken, how often the run ended degraded. Those three move before anyone notices a quality problem, and they need no evaluation harness.
A retention window you chose on purpose. Request and audit records are kept for 7 days on the entry plan, then 30, 90 and 180 as the plans grow — on the site's current plan table those are the four numbers a chain has to plan around. Longer-horizon evidence is the operator's to store, and a compliance requirement measured in years is a storage decision, not a platform setting.
multi agent framework: when a framework is the wrong first answer
The phrase describes both a category of libraries and a design choice, and the choice is the one that matters. A framework earns its place when the topology is genuinely complex: many interchangeable workers, a scheduler, retries with backoff across dependencies, state that outlives a single process. That is real work and rewriting it badly is a poor use of a quarter.
Three cases where it is the wrong first answer, all common in the first month of a project:
The graph is a line. If the steps run in a known order, a list and a loop is the whole implementation; a framework adds a second vocabulary for the same thing.
The state fits in one context. Multi-agent machinery exists to move state between contexts. If the state fits in one, the machinery is moving nothing.
Nobody can name the boundary. The boundary is the justification — two phases whose contexts make each other worse. If the justification is "it seemed more scalable", the second agent is a cost with no counterpart, and the pipeline would have been smaller and easier to explain.
The test worth applying before adopting anything: write the two prompts. If they read well as one prompt with two headings, it was one step.
research automation: the template that shows the shape
Research is the pipeline everyone builds first, and it is a good teaching example because its steps
are obvious: find candidates, read them, remove what overlaps, compress to fit, then write. Called
by name on the product side — research — that is the same chain the cluster centre describes,
which is the point: the shape recurs because the steps are the work.
Two lessons from it generalise.
The compression step is where pipelines are won or lost. Everything upstream produces text; the compression step decides how much of it the model ever sees. It is also the step most likely to degrade quietly, which is why its status has to be readable, and why a chain that reports compression honestly is easier to tune than one that reports a number.
The write step is where the design shows. A research chain that ends in "and now store the finding" has one non-idempotent step, one place where a retry duplicates, and one place where the key has to be a fact about the work rather than an index. Chains that get this right can be re-run freely; chains that do not get slower and stranger every time they are retried.
The first step is worth reading as code, because it is where the "one artefact out" rule is either honoured or lost. A search step has to sit in front of several engines, each with its own response shape, and its job is to make that difference disappear before the next step sees anything:
# backend/smartgate/modules/search/algorithm.py — source lines 102–118 (python)
class Search:
"""搜索容器 — SearXNG JSON / Firecrawl API / DuckDuckGo Lite。"""
effective_backend: str = ""
def __init__(
self,
search_query: SearchQuery,
settings: SearchRuntimeSettings,
):
self.search_query = search_query
self.settings = settings
self.result_container = ResultContainer()
self.start_time = None
self.actual_timeout = None
self.effective_backend = ""
# backend/smartgate/modules/search/algorithm.py — source lines 206–221 (python)
def _normalize_hit(self, item: dict, engine: str) -> dict:
snippet = (
item.get("markdown")
or item.get("description")
or item.get("snippet")
or item.get("content")
or ""
)
return {
"title": (item.get("title") or "").strip(),
"url": (item.get("url") or "").strip(),
"content": snippet.strip()
if isinstance(snippet, str)
else str(snippet),
"engine": engine,
}
Two things are happening in those two windows. The container carries the query, the runtime settings
and the result container, and it holds a field for which backend answered — that field is the
step's own provenance, kept beside the results rather than reconstructed afterwards. The normaliser
below it is the more important half: whether the backend returned markdown, description,
snippet or content, the hit that leaves this step has exactly four fields — title, url, content
and the engine that produced it. A downstream step therefore never learns which engine answered,
which is what lets the rest of the chain stay unchanged when a backend is swapped or a fourth one is
added.
That is the whole discipline of a pipeline step in one screen: accept the messy world, emit one shape, and record where the shape came from.
The comparison, in one table
| Shape | Order decided by | Cost profile | Failure diagnosis | Use when |
|---|---|---|---|---|
| Script / scheduled job | the author, statically | flat per run | logs and exit codes | the inputs have a known shape |
| Agent pipeline | the author, with named decision steps | per step, plus per decision | per-step offset and status | the work decomposes and the order is known |
| Loop (one agent, many turns) | the model, every turn | unbounded until the stopping rule | transcript only | the next action is genuinely unknown |
| Multi-agent system | a router or a supervisor | per hop, on top of per step | topology plus per-agent records | two phases need separate contexts |
The column that decides most projects is the last one, and it is answered by the work rather than by ambition. Reach for the row that matches your decomposition, not the row that matches the demo.
Frequently Asked Questions
Limitations
- A pipeline cannot absorb a new shape of input. Its strength is that the order is known; that is also its limit. Work that changes character mid-run needs a decision step, and a decision step is more expensive and less testable than the steps around it.
- Degradation has to be reported, not assumed. A fallback that passes text through unchanged is a reasonable outcome and a dangerous silence — a consumer that cannot tell it apart from a real compression will treat a degraded artefact as a result.
- Three of the four records are cheap; the fourth is not. Timing, status and input provenance come from the runner. Attributing spend per key and per team needs the metering layer, and a pipeline without it can tell you a step was slow but not what it cost.
- Adopting a framework is a topology decision, not a maturity signal. The vocabulary it brings is only worth it when the topology needs it.
- This page quotes no code, and that is a measured finding rather than a shortcut: see the method note. The product behaviour described above was read from the product source, not summarised from a fence.
- The demand numbers are a snapshot. One country, one window, one endpoint; the difficulty of the head phrase being 0 is a statement about competition for it today, not a guarantee about next quarter.
Sources
- The Model Context Protocol specification — introduction and the 2025-06-18 revision, for how a tool call is described and returned.
- Anthropic's engineering note on workflow and agent patterns — Building effective agents.
- Demand figures: this project's own paid measurement (DataForSEO Google Ads, United States,
12-month window, 2026-09-29), recorded in
search_volume.json. - Product behaviour: read from the product source at the revision pinned in this project's
pipeline_results.json; plan figures re-verified against/pricingon 2026-09-29.
Slice provenance
| # | SERP keyword | Symbol | File | Source lines | How it was pinned | sha256(12) |
| 1 | research automation | Search | backend/smartgate/modules/search/algorithm.py | 102–118, 206–221 | rule A L2 → slot-proof | 5d88cd6aed7c |
Method note
The matcher pinned 3 of 7 sections for this page (0 abstention(s), 4 no-slice verdict(s)), and one excerpt is quoted above. The other two "pinned" sections resolved to the same symbol — a Redis pipelining helper matched because the keyword contains the word "pipeline" — which is a generic-name collision rather than evidence about agent pipelines, so those two are written from sources like the rest. That is why this page carries one excerpt and not three.
Every product claim was read from the product source at the revision pinned in this project's
pipeline_results.json, read-only; the plan figures were re-verified against the live pricing page
on 2026-09-29. The quoted block is cut from the slice body and re-asserted against it byte for byte
before publication, and no batch fingerprints, auction data or internal hosts appear in the text.