Agentic Workflow Examples: Six Patterns You Can Run
A usable agentic workflow example states four things: the input it consumes, the steps it runs in order, what happens when a step fails, and one acceptance signal that says the run worked. The six patterns below cover the shapes that keep appearing in production — research to decision, nightly batch enrichment, a review loop, fan-out and merge, routing with escalation, and a monitoring watch —…
Short answer: A usable agentic workflow example states four things: the input it consumes, the steps it runs in order, what happens when a step fails, and one acceptance signal that says the run worked. The six patterns below cover the shapes that keep appearing in production — research to decision, nightly batch enrichment, a review loop, fan-out and merge, routing with escalation, and a monitoring watch — and each is written as those four fields rather than as a diagram.
Key takeaways
- The examples are ordered by how much the model is allowed to decide, from a fixed sequence to a routed judgement call.
- Each one needs the same three controls: a cap on steps, a written rule for a failed step, and one countable acceptance signal.
- Start with the batch example if nothing agentic is in production yet — its failures are the most visible and its lessons are the cheapest.
- Read the acceptance signal as a cost control: it is the number that tells you whether the next run is worth paying for.
Agentic workflows examples, and what makes one runnable
Most published examples are diagrams. A diagram of boxes and arrows is a drawing of a shape; it is not something you can run, and it will not tell you what to do on the Tuesday when a retrieval step returns nothing and the model cheerfully continues without it. The six patterns on this page are written to be run, and that means each one states four fields:
- Input — what the run consumes, in enough detail to recognise it in your own systems.
- Steps — the ordered work, with the point where the model is allowed to decide marked.
- Failure handling — what happens at each step that can fail, including the step that fails by returning nothing.
- Acceptance signal — the one measurement that says this run was good, and that a later run can be compared against.
The fourth field is the one that gets skipped, and it is the one that decides whether the example survives contact with a real workload. An example set without acceptance signals teaches a reader to ship something and hope. "The answer looked good" is not a signal; "every claim in the output names the source it came from, and the run finished inside eight tool calls" is.
Run-ability is also why none of these examples names a framework at the top. The control flow is the transferable part, and the same four fields describe a hand-rolled loop, a durable workflow engine, or a hosted builder. The vocabulary is not the subject here: whether the thing is best called a loop, a pipeline or a multi-agent system is the map kept on the agentic workflow vocabulary page, and this page assumes that map and starts from the shapes people actually put behind a scheduler.
One further discipline runs through all six: the example must state what it refuses to do. A research workflow that cannot say "the evidence is insufficient" will invent an answer, and a routing workflow that cannot hand work to a person will guess. The refusal is part of the design, not an error path bolted on afterwards.
Agentic workflows AI: the three places the model actually decides
Before the examples, it is worth naming how little of each one is actually model-driven. Across all six patterns the model makes at most three kinds of decision, and being explicit about which ones are in play is what keeps a run explainable after the fact.
Routing — which step runs next. This is a decision about the task: which queue, which specialist, which of two branches. It is the cheapest decision to constrain, because the candidate set is yours; a closed list of six routes is a classification problem, and an open-ended "decide what to do" is how a run spends an hour and a budget discovering it should have asked.
Interpretation — what a step's output means to the next step. Extracting a date, a customer name, a risk level or a set of entities from unstructured text is where a model earns its cost, because the alternative is a brittle regular expression. Interpretation is also where output contracts matter most: a typed result the next step can validate is worth more than a fluent paragraph it has to read.
Judgement — whether the task is done, and whether the result is good enough to release. This is the decision to keep closest to a human. An agent that grades its own work will pass itself; the useful version of judgement is a checkable property (a schema, a citation, a test, a required refusal) plus a person for the residue.
Everything else in these patterns should be deterministic: collection, chunking, retries, offsets, deduplication, writes, and the record. The dividing line matters because each of the three decisions above needs a different control — routing needs a closed candidate set, interpretation needs an output schema, judgement needs a check and an escalation path. Which classes of tool own those controls, and which of them are worth buying rather than writing, is the subject of the agentic workflow tool stack; here they appear only as the parts each example assumes exist.
Example 1 — the agentic workflow builder: research to a decision
The first pattern is the one most teams describe when they say "agentic": a question goes in, a recommendation comes out, and several sources were read on the way. It is also the pattern that wastes the most money, because the natural implementation loops until the model feels finished.
Input. A decision question with a deadline, a set of sources (internal documents, a public corpus, a search index), and a decision template with fixed fields — recommendation, the two strongest arguments against it, the evidence for each, and what would change the answer.
Steps. Scope the question into three to six sub-questions that can each be answered from one source at a time. Run one retrieval step per sub-question. Synthesise the retrieved passages into the decision template. Stop.
Failure handling. A retrieval step that returns nothing returns "insufficient evidence" for that sub-question — it does not return a plausible paragraph. A synthesis that cannot attach a source to a claim drops the claim rather than keeping it unsourced. The step counter is capped (in practice eight tool calls is generous for this shape), and hitting the cap ends the run with a partial answer that is labelled partial.
Acceptance signal. Every claim in the output names the source it came from, the run finishes inside the step cap, and the recommendation is one of the values your template allows. If a reader has to ask where a sentence came from, the run failed regardless of how good the sentence reads.
The economics are the reason to be strict here. A research run is the shape where a missing stop condition costs the most, because each extra loop re-reads material the previous loop already paid for. Writing the plan down as data — a list of sub-questions, each with a status and a source — is what makes the run inspectable afterwards, and that list is the "builder" part of the pattern. No product is required to hold it; a table is enough.
Example 2 — the agent pipeline: nightly batch enrichment
The second pattern is the one with the least glory and the best return: a list of records goes in, the same record comes back enriched, and it runs while nobody is watching.
Input. A list of records (anywhere from a few hundred to a few hundred thousand), one output schema, and a per-record instruction. Also, in practice, a cost ceiling for the batch.
Steps. A scheduler splits the list into chunks of a few hundred records. Each chunk becomes one task carrying an idempotency key (the batch id plus the chunk offset). A worker calls the model per record with the output schema enforced, validates the result against the schema, and writes it keyed by the record id. The offset of the last committed chunk is recorded so a restarted run resumes rather than restarts.
Failure handling. A record whose output fails validation is retried once with the same key; a second failure sends it to a dead-letter list instead of stopping the batch. A worker crash leaves the offset where it was, so the next run re-processes one chunk — which is safe precisely because the write is keyed. The dead-letter list is reviewed by a person, because a record that fails twice is usually a schema problem rather than a model problem.
Acceptance signal. The number of written rows equals the number of input records minus the dead-letter list, and re-running the same batch writes zero new rows. That second number is the real test: it is the difference between an idempotent pipeline and one that quietly double-charges every time a worker is recycled. The chain mechanics underneath — what a step receives, where the time went, what a retry costs — are worked through on the agent pipeline walkthrough.
Example 3 — orchestration frameworks in a review loop
The third pattern produces a patch instead of an answer: something is drafted, something reviews it, and the pair iterate until the result passes checks or a person takes over.
Input. A draft (a document, a pull request, a prompt, a configuration) plus a rubric of checkable properties — the things that must be true of the output, written as assertions wherever possible.
Steps. A reviewer step flags findings with their locations. A revision step produces a candidate patch for those findings. A deterministic check (schema validation, lint, a test) runs against the candidate. The loop repeats while findings remain and the round counter is under its cap.
Failure handling. The round cap is two or three, not ten; at the cap the run stops and hands the diff so far to a person with the outstanding findings attached. When the reviewer and the deterministic check disagree, the deterministic check wins — a model's objection to a passing test is an opinion, and treating it as a blocker is how a loop never terminates. A patch that fails the check is labelled "needs human" rather than re-prompted indefinitely.
Acceptance signal. Re-running the same input produces the same number of findings, the loop stops at the cap rather than before it, and every patch is either passing the deterministic check or explicitly labelled for a person. That first number is what makes the review reproducible; a reviewer whose finding count moves between identical runs is a source of noise, and the rubric is usually the thing to fix. The semantics of who runs when, and what a handover has to carry, are on the orchestration semantics page.
Example 4 — multi agent systems examples: fan out, then merge
The fourth pattern is the only one where more than one worker is genuinely justified, and the justification is parallelism rather than a second opinion.
Input. One task that decomposes into independent sub-lookups — the same question asked of six regions, six repositories, six suppliers — plus a merge contract describing what the combined result must contain.
Steps. A supervisor splits the task into N sub-tasks, binds a per-worker budget to each, and dispatches them. Each worker returns a typed result: status, payload, and the source it used. The merge step deduplicates and reconciles, then writes one result with a coverage field describing how many of the N sub-tasks reported.
Failure handling. A worker that times out returns a partial result with its status set, and the merge tolerates a missing set — a merged answer covering five of six regions is useful and honest, while a merged answer pretending to cover six is a defect. Conflicts between workers are resolved by source recency, or flagged rather than averaged. Because each worker is a separate task with its own key, a single failed worker can be re-run alone, which is the operational reason to split at all.
Acceptance signal. The merged result always carries its coverage count, the run's spend stays inside the supervisor's budget multiplied by N, and re-running one worker changes exactly one part of the output. What survives between the workers — which is nothing unless you decide otherwise — is a memory decision, and that decision is worked through on the agent memory architecture page.
The anti-pattern to avoid is the "team" whose members all read the same context and call the same tools. It costs a multiple of one agent to produce a reworded version of one answer, and its failure mode is invisible: the bill triples while the output looks roughly the same, which is why the coverage field above is worth more than any quality claim about the merge.
Example 5 — an agentic workflow that routes, then escalates
The fifth pattern sits in front of a queue and decides what happens to each item. It is the pattern with the clearest business case and the highest blast radius, because an automated action is a commitment.
Input. An inbound item (a support ticket, a form submission, an inbound message) plus a routing table and an escalation policy. Both are data, not prompt text, so they can be reviewed and changed without a deploy.
Steps. Classify the item into a fixed set of intents and a risk level. Route by the table: low-risk intents to an automated response, medium to a human review queue with a suggested reply, high-risk straight to a person with no automated action. A confidence threshold gates the automated branch.
Failure handling. Below the threshold, the item goes to a person — the automated branch fails closed, never open. High-risk categories never reach the automated branch even at maximum confidence. If the same input classifies differently on two consecutive runs, the item is routed to a person and the classifier is treated as faulty: a router that flip-flops is a bug, not a judgement call.
Acceptance signal. Two rates, measured continuously: the share of items resolved without a person, and the escalation rate. Both are only trustworthy alongside a sampled review of the automated replies, because a router that escalates nothing looks efficient until someone reads what it sent. The refusal list in the first section is what keeps this pattern honest — the workflow must be able to say "not mine" and hand the item on.
Example 6 — AI workflow automation for a monitoring watch
The sixth pattern runs on a schedule and mostly does nothing, which is exactly what you want from it. It watches a handful of signals, and it speaks only when something is material.
Input. A schedule, a set of watches (quota usage against a limit, an error rate, a search term, a competitor's release feed, a document that should have changed), and a definition of material.
Steps. Collect the readings, read-only. Compare against the previous reading and against the threshold. Classify whether the change is material. For a material change, draft a short note and open one queue item carrying the reading, its timestamp and the comparison it was made against.
Failure handling. A watch whose collection fails reports "no data", never "no issue" — the two are different states and conflating them turns an outage into silence. A dedupe key per watch per window prevents one event from opening twenty queue items. The whole workflow runs read-only by default; anything that writes, closes or notifies is a separate, explicitly enabled action.
Acceptance signal. Every raised item carries the reading and its timestamp, the false-positive rate is measured per watch class rather than in aggregate, and a dry-run mode exists that produces the notes without writing them. The last one is not a nicety: it is how you tune a threshold without mailing twenty people about a number that turned out to be normal. Which processes deserve this machinery at all, and what the judgement step costs when you make it explicit, is the adoption question on AI workflow automation.
The six examples, in one table
The table is the shortest useful summary of the page: for each pattern, what goes in, the failure you must handle before anything else, and the single number that says the run worked.
| Example | Input | The failure you must handle first | Acceptance signal |
|---|---|---|---|
| 1 · Research to a decision | A question, sources, a decision template | A retrieval step that returns nothing and a synthesis that continues anyway | Every claim names a source; the run ends inside the step cap |
| 2 · Nightly batch enrichment | Records, an output schema, a cost ceiling | A record that fails validation twice, and a worker restart | Written rows equal input rows minus dead letters; a re-run writes nothing |
| 3 · Review loop | A draft and a rubric of checkable properties | The loop that never terminates while the reviewer keeps finding something | The same input yields the same finding count; the loop stops at the round cap |
| 4 · Fan-out, then merge | One task, N independent sub-tasks, a merge contract | A worker that times out and a merge that pretends to be complete | The result carries its coverage count; one worker can be re-run alone |
| 5 · Route, then escalate | Inbound items, a routing table, an escalation policy | The confidence threshold that fails open instead of closed | Auto-resolution and escalation rates, plus a sampled review of automated replies |
| 6 · Monitoring watch | A schedule, a set of watches, a definition of material | A failed collection reported as "no issue" instead of "no data" | Every item carries its reading and timestamp; false positives measured per class |
Read the middle column as a sequence when you are choosing. The first failure in each row is the one that causes real damage rather than visible breakage, and it is the one an example page usually omits: an empty retrieval is invisible, a silent watch looks calm, and a merge with missing coverage reads like a confident answer.
Where SmartGate fits
Every example above assumes the same two things: one place that sees the model call, and one record of what it cost. SmartGate is the gateway in that position — each tool call through it is written as an audit row (caller, route, transport, tool, token count, latency, outcome), and that same row is what the per-key rate limit and the per-team budget read. For the batch and fan-out examples the budget is the subject; for the routing and monitoring examples the audit row is the evidence you attach to a decision.
The operational limits come from the plan rather than from the feature list: monthly token caps of
2M, 20M, 100M and 200M+, requests per minute per key of 120, 300, 600 and 1200, audit-log retention
of 7, 30, 90 or 180 days, and 2, 10, 30 or unlimited keys per team. If a retention window or a rate
limit is what a workload depends on, compare it against the current pricing table on /pricing
before committing a date, because those numbers move with the plan and the published table is the
authority rather than this page.
Where a workflow needs a refusal it cannot express in a prompt — a hard cap, a scoped credential, a per-team ceiling — that belongs at the gateway rather than in the instructions. The examples are written expecting that: their failure handling assumes something outside the model can say no.
Frequently Asked Questions
Limitations
These are shapes, not benchmarks. No example is a claim that a particular framework, model or vendor implements it well, and no timing or cost figure is quoted, because both depend on the workload rather than on the pattern.
The six patterns do not cover every production workflow. They deliberately omit long-running human-in-the-loop processes, anything that writes to a system of record without review, and the fine-tuning and evaluation pipelines that sit beside an agent stack rather than inside it.
The failure handling described here is the minimum, not a full account. A real deployment also needs rate limiting against upstream APIs, secret rotation, and a runbook for the case where the record itself is unavailable — none of which is specific to agentic workflows, and none of which is covered above.
This page carries no code excerpt, and the reason is recorded in the Method note below: the slice matcher found no unique symbol for any of its eight section keywords. The consequence is honest: no line-numbered claims, no quoted implementation detail, and no assertion about the internal shape of any product named above.
Sources
- Anthropic's engineering note on workflow and agent patterns — Building effective agents, for the workflow-versus-agent split and the orchestrator and evaluator shapes the review loop follows.
- The Model Context Protocol specification — modelcontextprotocol.io/specification, for how a runtime tool is described and what a client may assume about it.
- Temporal's documentation — docs.temporal.io, for durable execution, idempotent steps and resuming from a recorded offset, which the batch example depends on.
- OpenTelemetry's semantic conventions for generative AI — opentelemetry.io/docs/specs/semconv/gen-ai, for the span attributes that make a run's cost and steps readable afterwards.
- Google's SRE Workbook — sre.google/workbook/table-of-contents, for the alerting and escalation discipline the monitoring and routing examples borrow.
- Demand figures on this page are our own measurements: DataForSEO Google Ads, United States,
12-month window, measured 2026-09-30, recorded in this project's
search_volume.jsonandresearch_brief.md. - Product behaviour and the plan table: read read-only from the product source at the revision
pinned in this project's
pipeline_results.json, with the plan figures re-verified against the live pricing page on 2026-09-30.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice
matcher pinned 1 of 8 sections for this page (0 abstention(s),
7 no-slice verdict(s)). The single pin is the same generic collision the rest of this
cluster recorded: a Redis pipeline helper whose name contains the word "pipeline", matched because
a section keyword contains that word rather than because the symbol is an agent pipeline. A block
fenced from that helper would have dressed the page as verified work, so every section above is
written from sources instead.
Product claims were read from the product source at the revision the slice run recorded in this
project's pipeline_results.json, read-only, and the plan figures were re-verified against the live
pricing page on 2026-09-30. The section keywords quoted above each heading come from this project's
own paid measurement run rather than from a third-party tool. No code, batch fingerprints, auction
data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.