SmartGate

What Is an Agentic Workflow? Loop, Pipeline and Orchestration

An agentic workflow is a workflow whose next step is decided at run time — usually by a model choosing a tool — instead of by a fixed graph. It needs four things to be operable: steps, the state carried between them, a stopping rule, and a record of what actually ran. The phrase carries roughly 3,600 US searches a month, so the definition is worth getting right.

Short answer: An agentic workflow is a workflow whose next step is decided at run time — usually by a model choosing a tool — instead of by a fixed graph. It needs four things to be operable: steps, the state carried between them, a stopping rule, and a record of what actually ran. The phrase carries roughly 3,600 US searches a month, so the definition is worth getting right.

Key takeaways

  • The decision point is the dividing line. If a human drew the graph in advance, it is automation; if the model picks the next tool at run time, it is agentic — and it inherits the failure modes of a decision instead of the failure modes of a schedule.
  • A workflow without a stopping rule is a billing event. Loops need a step budget, a spend cap and a retry policy that names which of the two it hit.
  • Steps should be boring and observable. What makes a workflow debuggable is not the model, it is whether each step records its own timing, its input and its output.
  • Multi-agent is a cost multiplier before it is a capability multiplier. Supervision and handoff buy separation of context; you pay for it in tokens and in latency on every hop.
  • The three names engineers search for — orchestration, pipeline, workflow — describe the same object from three angles. Orchestration is who decides, pipeline is what runs in order, workflow is the whole thing plus its state.
  • Write the stopping rule and the record before you write the prompt. Start with the four questions in the first section: what are the steps, what state moves, when does it stop, and where is the evidence.

What an agentic workflow is, and what it is not

An agentic workflow is a control loop with tools. The loop does three things: look at the current state, decide the next action, execute it and fold the result back in. The decision is the part that is not scripted. Everything else — the tool call, the parsing, the retry — is ordinary software that happens to be inside a loop.

That is the difference from an automation you would build with a visual builder. A scheduling graph has edges an author drew; when the environment changes shape, the graph keeps walking the same edges and produces a plausible-looking wrong answer. An agentic loop re-decides at every turn, which is why it absorbs variation and why it can wander. Both behaviours come from the same property.

Three consequences follow, and they are the reason the rest of this page is organised the way it is:

  1. The step vocabulary has to be closed. A model that can call any of forty tools will call the wrong one; a model that can call seven will call the wrong one less often. Closing the vocabulary is a design decision made before the first prompt.
  2. State is a first-class object. In a drawn graph the edge carries the data. In a loop the context window carries it, which means it grows, which means it has to be managed.
  3. Cost is a function of turns, not of pages. A fixed job that touches 100 records costs what the 100 records cost. A loop costs whatever it costs to reach its stopping rule, and the stopping rule is the only thing standing between you and a runaway.

If you take one sentence away: an agentic workflow is the part of the system that decides, and deciding is the part you have to instrument.

ai workflow automation: automating steps versus automating judgement

The phrase people search for and the thing they mean are one step apart, and the gap is where most failed projects sit. Automation replaces a manual step with a deterministic one: fetch the invoice, parse the total, write the row. It is reliable in the way a script is reliable — it either works on every input of the shape it expects, or it fails loudly on the ones it does not expect.

Judgement automation is different. Deciding which of three documents is the relevant one, whether a search result answers the question, or whether to keep digging — those are calls a person makes, and they do not have a schema. You cannot write them as a step because you cannot enumerate the input space in advance. This is the honest boundary between ai workflow automation and an agentic loop: the first automates the steps around a judgement, the second automates the judgement itself and keeps the steps.

The practical arrangement that survives contact with production is a pipeline that contains one or two judgement points rather than a loop that contains everything. Put the deterministic work in steps: fetching, extracting, deduplicating, compressing, writing to memory. Put the model at the seams where the next action is genuinely unknown. You get the auditability of a pipeline for the 90% that is mechanical, and you contain the part that is not.

That split is also what makes a workflow measurable. If a run consists of six steps and one decision, a bad outcome has seven candidate causes; if it consists of one model call, it has one cause you cannot decompose. Teams that skip this step end up rewriting the prompt when the real fault was a fetch that returned an empty body — which is exactly why the sections below are about timing, retries and records rather than about prompting.

agent orchestration: who decides the next step

Orchestration is the answer to a single question: at each turn, what decides what happens next? There are three answers, and choosing among them is the main architectural decision in an agent system.

Code decides. The order is fixed and the model is called with a specific prompt at each stage. This is a pipeline with model calls inside it. It is the cheapest to operate, the easiest to test and the least interesting to demo.

The model decides. The loop runs until the model says it is done. This handles open-ended tasks and is where the cost variance lives: the same input can take three turns or thirty.

Something in between decides. A router classifies the request and hands it to one of a few sub-workflows; each of those is a fixed chain. Most production systems that call themselves agentic are this third shape, because it bounds the variation while keeping the flexibility where it pays.

Two rules make any of these operable. First, an orchestrator is not a step: a pipeline tool that can execute other tools must refuse to execute itself, or the first recursive call becomes a stack overflow that bills until someone notices. Second, the scorer is code: the thing that decides "is this good enough to stop" should be a deterministic check wherever a deterministic check exists, because a model grading its own output is a loop with no floor.

Where orchestration becomes a cost problem rather than a design problem is discussed in traffic shaping and rate limits — a fleet of loops hitting one upstream has a shape, and it is not the shape a single interactive user has.

agent pipeline: stages, state and what each step returns

Once the order is fixed, a pipeline is five decisions: what the steps are, what each one receives, what each one returns, what happens when one fails, and what is recorded. The interesting one is the second, because that is where pipelines usually leak.

A step should receive exactly what it needs, from the previous step, by name. When a pipeline passes an accumulating blob of everything so far, three failures become invisible: a step that silently ignores its input, a step that re-does work an earlier step already paid for, and a context window that grows until compression stops being a choice.

Concretely, a well-behaved step contract has four fields worth insisting on:

Field What it answers Why it matters at 3 a.m.
input which prior step's output this step consumed a wrong answer with a stale input is a wiring bug, not a model bug
output the artefact, in the shape the next step expects lets you re-run one step instead of the whole chain
timing how long this step took, and how far into the run it started separates "the model is slow" from "the fetch is slow"
status success, degraded, or failed, plus which branch was taken the difference between a wrong result and a missing one

That is not a theoretical list. It is the difference between a pipeline you can operate and a pipeline you can only re-run.

pipeline orchestration: retries, offsets and idempotency

Pipeline orchestration is the unglamorous half: what happens on the second attempt. Three properties decide whether a retry is safe.

Idempotency. A step that writes must be safe to run twice. Fetching and compressing are naturally idempotent; "add this to memory" is not, which is why the write steps are the ones that need an explicit key rather than a position in a sequence.

Offsets, not just totals. A total runtime tells you the run was slow. A per-step offset tells you which step was slow, and whether the slowness was the step or the wait in front of it. Recording the offset of each step — when it started, how long it took, how much of the run it owned — is the cheapest observability you will ever buy, and it is the field most pipelines omit.

A named branch on failure. When a fallback fires, the result should say so. A degraded result that is indistinguishable from a good one is worse than an error, because the error at least stops the chain. The pattern worth copying is a status field that says degraded with the reason, so that a downstream consumer can decide whether a 1.0x compression ratio means "nothing to compress" or "compression gave up and we passed the original text through".

Those three properties are what let a pipeline survive a bad day without a human reading the logs of every run. They are also what makes the next section possible at all: you cannot add an agent to a pipeline that cannot tell you which step failed.

agent supervisor: the pattern, its bill, and when it pays

The supervisor pattern puts one agent in charge of several. The supervisor holds the task, decides which worker to invoke, inspects the result and decides whether to invoke another. Workers do not talk to each other; they receive a scoped instruction and return an artefact.

It pays in exactly one situation: the sub-tasks need different context and the contexts conflict. A research task that must read ten documents and write one summary does not need a supervisor — it needs a fetch step, a dedup step and one writer, because all ten documents belong in the same context. A task that must apply a legal policy and also produce a price quote does, because the two rule sets pull the context in different directions and mixing them degrades both.

The bill is structural rather than incidental. Every hop through the supervisor is another model call that carries the supervisor's context, so the supervisor's context window becomes the busiest object in the system, and its cost grows with the number of workers it manages rather than with the work. A supervisor with six workers can spend more on deciding than all six workers spend on doing.

The mitigation is the same in every system that does this well: the supervisor's job is to route and to check, not to re-read the workers' inputs. It sees the instruction it sent and the artefact it got back, not the ten documents underneath. Keeping that boundary is what keeps the pattern from collapsing into "one agent with a very large prompt".

agent handoff: passing a task without losing the thread

A handoff is the moment one agent stops and another continues. It is not the same thing as a supervisor delegation, and confusing the two is the most common modelling error in multi-agent design: a handoff transfers ownership, a delegation keeps it.

What has to survive a handoff is small and specific. The goal, restated in the receiving agent's vocabulary. The state that has been established and must not be re-derived — which URLs were already read, which were dead, which were deduplicated. The limits: what the receiver may spend, how many turns it has, and what it must not touch. And the answer to "how do I know I am done", because a receiver that inherits a goal but not a stopping rule will keep working.

Two failure modes are worth naming because neither looks like a failure at the time. The first is the silent restart: the receiver does not receive the established state, re-fetches what the previous agent already fetched, and the run costs double while looking normal. The second is the inherited assumption: the previous agent's guess becomes the next agent's fact, and no single step in the chain is wrong enough to be caught.

The fix for both is to make the handoff an artefact rather than a conversation. A structured payload — goal, established state, limits, done-condition — is checkable before the next agent starts, and it is the object you can attach to the record when you need to explain what happened.

multi agent orchestration: three topologies and their failure modes

Multi-agent orchestration is a topology choice, and there are three that actually get used.

Topology Shape Fails by Use when
Sequential pipeline a fixed chain of specialised steps one step's bad output becomes the next step's input the work decomposes cleanly and the order is known
Supervisor one router, several workers the supervisor's context becomes the most expensive object in the system sub-tasks need conflicting contexts
Handoff chain ownership moves between peers state does not survive the move and work is repeated the task changes character partway, e.g. discovery then execution

What none of the three fixes is the multiplicative cost. A topology with three agents in series makes three model calls per turn where one made one; a supervisor adds a routing call on top. The phrase carries 720 measured searches a month with a difficulty of 22, and the sensible reading of that demand is that people are looking for the topology decision, not for permission to add agents.

The rule that keeps this honest: add an agent when there is a context boundary to draw, not when there is work to do. Work parallelises with a queue; context does not.

multi agent ai: when one agent is the better answer

A single agent with four well-named tools beats a multi-agent system with the same four tools in almost every case that fits in one context window. That is not a conservative position, it is the arithmetic: one loop has one stopping rule, one cost curve and one place to look when it goes wrong; a system of three loops has three of each, plus the interface between them.

The cases where multi-agent genuinely wins share a property: the contexts are not just different in size, they are different in kind, and letting them mix makes both worse. Examples that hold up: a retrieval-heavy phase followed by a compliance-heavy phase; a long-running research phase that would fill a context window with material the execution phase must not see; a write phase that needs a small strict context after a read phase that needed a large loose one.

Everything else is usually a pipeline with a bigger prompt. If you cannot name the boundary, you do not have one — and the test is cheap: describe the two prompts. If they would read well as one prompt with two headings, they are one agent.

llmops: how you operate a workflow after it ships

Operating a workflow is a different job from building it, and it is where the budget actually goes. Four things have to be visible from outside the loop: what ran, what it cost, whether it worked, and what changed.

What ran is the audit record: one row per tool call, with the caller's identity, the arguments that matter, the outcome and the correlation id that groups the calls belonging to one task. This is the layer LLM observability describes field by field, and the reason it is worth insisting on is that a task assembled from five calls in three services is otherwise un-reconstructable.

What it cost has to be attributable per key and per team, not per site, or the first runaway loop reads as "the API got expensive this month". Aggregated spend cannot tell you which caller did it.

Whether it worked is an inventory of judgements, not a single score: how often the stopping rule fired late, how often a fallback branch was taken, how often the run ended with a degraded status. Those three counters catch most silent regressions without an evaluation harness.

What changed is retention. Retention windows differ by plan — SmartGate keeps request and audit records for 7 days on Free, 30 on Pro, 90 on Teams and 180 on Enterprise — and a longer window is what makes a month-over-month comparison possible at all. If your compliance story needs a year, that part is yours to store, and it is better to know that before an audit than during one.

The operational surface a workflow's own bill lives on is token control; the traffic and audit pillars are described on their own pages, and neither replaces the record described above.

The three named templates, and what a step is

A pipeline needs a starting shape, and the most common starting shapes are the same three tasks in every codebase. SmartGate's smart_pipe tool ships them as named templates, so a caller gets a working chain without describing it:

Template Steps, in order What it is for
research search (5 results) then fetch, dedup at 0.85 similarity, then context compression at ratio 0.4 "answer this from the live web without stuffing the context window"
read fetch, then compress to a 2,000-token target "turn this URL into something an agent can hold"
remember search, then write the result into team memory "make this finding available to the next run"

A custom chain is a list of steps, each naming a tool and its parameters, optionally naming itself so that later steps can refer back to it, and optionally declaring which previous step's output it wants as its input. Two things about that contract are worth copying:

  • Step order is data, not code. The same runner executes a template and a hand-written chain, so a chain you can write in a config file is a chain you can change without a deploy.
  • An orchestrator cannot be a step. smart_pipe refuses to run itself as one of its own steps — it is the thing that runs steps, not a step. That single guard is what prevents the recursive pipeline that spends money in a loop nobody wrote.

What a workflow costs on each plan

Per-key rate limits are what a single loop meets; the plan is what its team meets. The four tiers differ on four numbers that matter to a workflow: monthly token allowance, requests per minute per key, log retention, and how many keys the work can be spread across.

Plan Monthly allowance Rate limit Log retention API keys Share of savings
Free 2M tokens 120 req/min/key 7 days up to 2 —
Pro 20M tokens 300 req/min/key 30 days up to 10 capped near 36 USD/month
Teams 100M token pool 600 req/min/key 90 days up to 30 capped near 100 USD/month
Enterprise 200M+ tokens, contract 1,200 req/min/key 180 days contract HMAC budgets

The pricing model is worth reading closely because it is unusual: you pay for the platform, and the savings share only starts once a team has saved 15 USD in a month, then stops at a cap. A workflow that compresses and deduplicates its own context therefore has a measurable effect on its own bill, and the effect is visible in the same place as the spend.

Where SmartGate fits

SmartGate is an MCP-native algorithm gateway — token control, traffic shaping and agent audit, driven by seven algorithm primitives. For a workflow, the relevant part is the middle of the chain: smart_fetch turns a URL into Markdown, smart_search aggregates engines, smart_context_gate compresses with a configurable ratio, smart_dedup removes overlapping passages, smart_budget_guard checks, counts and records spend with a hard ceiling, smart_memory stores team memory, and smart_pipe composes them. It is the store, not the model: it does not route between models, it does not tune itself, and it has no world knowledge of its own.

If you are building the workflow now, the sequence that avoids rework is: connect a client through the MCP gateway so every call is attributable from the first run, run the research template against one real question, then read the record it leaves. The audit and compliance surface is where those rows land.

Start from the pricing page to size the tier, or start with a key and run one template.

Frequently Asked Questions

Limitations

  • A fallback can be honest or invisible, and honesty costs a field. When context compression fails or returns nothing inside a template, the sensible behaviour is to pass the original text through rather than fail the run — but a passthrough that reports a 1.0x ratio must be distinguishable from compression that found nothing to remove. Any consumer that ignores that field will treat a degraded step as a good one.
  • Availability beats refusal in some paths. A rate-limit check that cannot reach its counter store has to choose between refusing every call in a fleet and allowing them; allowing, with the choice stated in the response, is the more common production answer — and it is a choice you should know your own system made.
  • Retention is per plan, not per ambition. 7 to 180 days of logs is enough to operate and not enough to audit a year. Long-horizon evidence is the operator's to store.
  • A gateway is not a model router and not an orchestrator of models. If your workflow needs model selection or automatic tuning, that lives in your application, not in this layer.
  • Multi-agent is not a capability you buy, it is a cost you choose. Every additional agent adds a model call per turn and a boundary that can lose state.
  • This page quotes no code. See the method note: the matcher's two hits were the same generic symbol, so nothing was quoted rather than quoting something that only shared a word with the subject.

Sources

  • The Model Context Protocol specification and introduction — modelcontextprotocol.io/introduction and the 2025-06-18 specification revision.
  • The reference MCP server collection — github.com/modelcontextprotocol/servers.
  • Anthropic's engineering note on workflow and agent patterns — Building effective agents.
  • Demand figures in this page are our own measurements: DataForSEO Google Ads, United States, 12-month window, measured 2026-09-28/29, recorded in this project's search_volume.json.
  • Product behaviour: read from the product source (backend/smartgate/core/pipeline.py, backend/smartgate/api/mcp_tool_docs.py) at the revision the slice run recorded in this project's pipeline_results.json; plan figures re-verified against /pricing on 2026-09-29.

Method note

This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher returned 2 of 10 sections pinned, 0 abstentions and 8 no-slice verdicts, all recorded in this project's pipeline run record. Both "pinned" sections resolved to the same symbol — pipeline in lib/redis/upstash-adapter.ts, a fourteen-line Redis pipelining helper matched because the keyword contains the word "pipeline" — which is a generic-name collision, not evidence about agent pipelines. Quoting it would have given the page the shape of a verified article with none of the substance, so every section above is written from sources instead.

Product claims were read from the product source at that same revision, read-only, and the plan figures were re-verified against the live pricing page on 2026-09-29. No batch fingerprints, auction data, internal hosts or code are transcribed: there is nothing quoted here to assert verbatim. The demand figure quoted per section comes from this project's own paid measurement run.