SmartGate

LLMOps: The Four Layers an Agent Stack Must Cover

LLMOps is the practice of running language-model systems as production software, and it has four layers — visibility, evaluation, orchestration and memory. Those are layers rather than products, and every tool in this field is selling you one of them.

Short answer: LLMOps is the practice of running language-model systems as production software, and it has four layers — visibility, evaluation, orchestration and memory. Those are layers rather than products, and every tool in this field is selling you one of them. This page is the selection guide: what each layer decides, which questions separate the options inside it, what is worth building yourself, and the order to add them so that no purchase is made twice.

llmops: the four layers, and what each one actually decides

The word covers a lot of ground, so it is worth separating what each layer answers. Visibility answers "what happened on this call" — the prompt size, the tool that ran, the tokens it cost, the latency, the error. Evaluation answers "did this change make the output better" — a set of cases with expected properties, run against two revisions. Orchestration answers "who runs next" — the loop, the branch, the retry, the stop condition. Memory answers "what survives this call" — the facts a later step is allowed to assume.

A stack that skips a layer does not fail loudly, it fails slowly. Without visibility, cost questions are answered by estimation and incident questions by guesswork. Without evaluation, every prompt change is a coin flip that ships on the author's confidence. Without orchestration, the control flow lives in a chain of callbacks nobody can read. Without memory, the same context is re-sent on every turn, which is the most common reason an agent bill grows faster than its usage.

The order matters more than the shopping list, and the cheapest ordering starts from the record. If one row per tool call exists — caller, route, tool, tokens, latency, outcome, and which task it belongs to — then evaluation, cost control and incident review all have something to read. That record is the subject of LLM observability, and it is the layer we recommend owning early, because the other three layers all consume it.

Where an agent stack sits behind a gateway, this layer has a second job: the record is also the enforcement point. A gateway that sees the tool call can apply a per-key limit and a per-team budget at the moment of the call, instead of after the invoice. That is the difference between observing spend and controlling it.

agent orchestration: who decides which step runs next

Orchestration is the part of the stack that decides the next action, and there are only three honest designs. A fixed pipeline runs a known sequence with retries and no branching; it is the easiest to reason about and the only one whose cost is predictable up front. A supervisor holds the plan and dispatches sub-tasks to workers, which is what people usually mean by an agent; the plan is data, so it can be inspected and replayed. A queue with handoffs moves a task between specialised agents, each owning one stage, which is the shape that scales operationally because a failed stage can be retried without restarting the task.

The mistake is treating orchestration as a framework choice first. The framework question only matters after two decisions are made: what the unit of work is (a tool call, a stage, a task) and where the state lives between steps. If the state lives in a durable store that a fresh worker can read, then the framework is replaceable and a crash is recoverable; if it lives in a process's memory, the framework is your reliability model, and upgrades become migrations.

Two operational properties separate real orchestration from a prompt chain. First, idempotency: a step that ran twice must not bill or write twice, which means each step carries a key and the store enforces it. Second, a stop condition that is not "the model said it was done": a maximum step count, a token budget, or a wall clock. Every runaway agent story is a missing stop condition. The vocabulary underneath all of this — what an agentic workflow is, and how the loop, pipeline and multi-agent shapes differ — belongs to this cluster's centre page. The mechanics of the chain itself (step contracts, ordering, what a run records) are worked through on agent pipeline; the decision of whether a process deserves this machinery at all is on AI workflow automation.

multi agent orchestration: supervisor, handoff, or a queue

Multi-agent orchestration is where budgets go to die, so it is worth being blunt about when a second agent earns its cost. There are three shapes and each has a specific reason to exist.

Supervisor and workers exist because the fan-out is genuinely parallel — one research task turns into six independent lookups whose results are merged. The cost model is simple: the number of workers multiplies token spend, so the supervisor must own a budget per worker, and the merge step must tolerate a partial result set.

Sequential handoff exists because the stages need different instructions, tools or permissions — a planning stage, a coding stage, a review stage. Here the reason to split is not parallelism but containment: the reviewer must not be able to write, and the coder must not be able to deploy. If two stages share the same tools and permissions, one agent with a longer instruction is cheaper and easier to debug.

Queue-driven orchestration exists because the work arrives continuously and needs to survive restarts: a task enters, a worker consumes it, the result goes back. This is ordinary backend engineering with a model inside it, and it is the shape that behaves best under load — each task has an independent timeout, retry and cost, and one bad task does not take the system down.

The orchestration anti-pattern is a "team" of agents that all read the same context and all can call the same tools; it costs a multiple of a single agent to produce a slightly reworded version of the same answer. A second agent is justified by a different input, a different permission, or a different output contract — never by the hope that another voice improves the prose.

multi agent systems: when more than one agent earns its cost

Multi agent systems sit at the top of this taxonomy, and the deciding question is not "how many agents" but "how many independent failure domains". Each additional agent adds three costs: the tokens it spends, the coordination surface it opens, and one more place for a non-deterministic result to enter the pipeline.

The pattern that survives production is few agents, explicit contracts. One supervisor that owns the plan and the budget; workers that receive a typed task and return a typed result; a single store that both read from. The pattern that does not is peer agents that negotiate, because negotiation is unbounded by construction and produces no artefact a reviewer can check.

When you do need many agents, the useful count is usually driven by a hard external constraint: a different tool set (a browser agent, a SQL agent, a deployment agent), a different trust level (untrusted input must not reach your write path), or a different latency budget (a fast extraction step in front of a slow reasoning step). Constraint-driven splits stay small and stable; capability-driven splits grow until someone re-draws the architecture — usually as a single agent with more tools.

ai agent framework: the four questions that decide it

The framework market is loud, so the useful move is to compare candidates on four questions that do not depend on branding.

  1. Where does state live between steps? A framework that persists state to a store you choose (Postgres, Redis, a queue) is recoverable and portable; one that keeps it in memory couples your reliability to a process. Ask for the exact table or key prefix the state lands in, and whether a second worker can resume a run.
  2. How is a step's cost attributed? You want the framework's own records to carry the model, token count and tool name per step, or you will be reconstructing cost from provider invoices. If the framework emits OpenTelemetry spans, the tracing layer can consume them without bespoke glue.
  3. What happens on a partial failure? Retry with the same key (idempotent), skip and continue with a degraded result, or fail the task. A framework that only offers the third option pushes the design problem into your application code.
  4. What does a stopped run look like? Being able to inspect a paused run — the plan so far, the step that failed, the tokens already spent — is what makes an agent operable at 3am.

Two practical filters cut the list fast. Frameworks that require you to adopt their hosting can be excluded if you need the record to stay in your own store; frameworks whose operators are only available in a UI can be excluded if you need the whole stack scriptable. The rest is taste and maintenance cost, and the honest tie-breaker is how readable the deterministic parts are: your engineers will spend more time in the retry and error branches than in the happy path. How a framework, a hand-rolled chain and a gateway layer divide that work is the comparison on AI Agent Architecture.

agent memory: what to persist, where, and for how long

Memory is the layer with the largest gap between what is demonstrated and what is operated. The distinction that matters is what kind of memory, because the three kinds have different storage, different retention and different failure modes.

Working context is what one task needs to finish: the current instruction, the tool results so far, the plan. It is short-lived by definition, and its cost is the reason context engineering is a discipline — every turn that re-sends the same documents is money spent on nothing. Semantic memory is the durable facts: which customer prefers which format, what the last decision was, which document is authoritative. It is retrieved by meaning rather than by key, so it needs an embedding index plus a decision about what is allowed to be written into it. Episodic memory is the record of what was done: the task, its steps, the outcome. That is the layer evaluation reads, and it is usually the same table as the observability record.

The operational questions are boring and decisive: what is the retention on each kind (7 / 30 / 90 / 180 days by plan, if you buy the gateway), what happens when retrieval returns something stale, and who is allowed to write. An agent that writes memory without review will happily persist a hallucination and then treat it as a fact on the next run — which is why the house rule is that semantic writes are a reviewed operation, not an autonomous one.

Where the context is the problem rather than the storage — long threads, repeated documents, large tool outputs — the cheaper fix is compression and de-duplication before storage grows. A token quota makes the same point from the other end: enforcing a token quota per team turns "our context is too big" from an opinion into a number, and token optimization techniques are the mechanical reductions that follow. What actually gets stored, when a read is allowed to run and what each path costs are worked through on Agent Memory Architecture.

ai agent tools: the categories, and what each one is for

The phrase "AI agent tools" covers two different things, and confusing them produces bad architecture. Tools the agent calls at runtime are its capabilities: a search, a database query, an API write, a file read. Tools you use to build and run the agent are your operational stack: tracing, evaluation harnesses, orchestrators, gateways.

For runtime tools, the design rules are the ones an API team already knows. Each tool needs a description precise enough for the model to choose it correctly, a typed input schema, a bounded output (truncate and summarise rather than returning a 200KB document into context), and — the rule that is easiest to skip — a permission boundary that lives outside the prompt. A tool the agent can call is a capability the agent has; describing it as "only for reads" in the system prompt is not a control. What that bound looks like once the material is already long — compression, segmentation, dedup and team memory before the call — is covered in Window Management Techniques.

For the operational stack, the categories are narrower than the market suggests. You need one record of calls (observability), one evaluation harness, one orchestrator, and one place where limits and credentials live (usually the gateway). Everything else is a feature of one of those four, and buying a fifth product to re-solve one of them is how stacks double in cost without getting better.

The MCP tools reference covers the runtime tool surface as it is described on the wire (schemas, annotations, what a client may assume), and the MCP gateway is where this page's operational layer is enforced in our own stack: one place that sees the tool call, applies the key's limit, and writes the audit row.

agent evaluation: what to measure before and after shipping

Evaluation is the layer teams postpone longest, because it looks like extra work before the first release and like insurance only after the first regression. The version that survives is small: a set of cases, each with a property that must hold, run automatically against a candidate revision.

Three properties make an evaluation set useful rather than decorative. It has cases, not prompts — an input plus something checkable in the output, whether that is a schema, a required fact, a refusal, or a tool call that must not happen. It has a baseline: the current revision's numbers, so the result is a delta rather than a score with no meaning. And it runs the whole pipeline, not the model in isolation, because most regressions arrive from a changed tool description, a new retrieval result, or a truncated context rather than from the model itself.

Where an evaluation cannot be automated, the honest fallback is a review queue over the real record: sample the tasks that used the most tokens, the ones that took the most steps, and the ones where a retry happened. Those three slices find most real defects, and they cost nothing beyond the record that the observability layer already keeps. The measurement discipline behind this — what a number from your own system does and does not tell you — is the same one behind the record layer described above — which is why observability and evaluation are usually built together in one release.

A buying checklist, in the order it matters

Written as a sequence, because each step makes the next one cheaper to decide.

Step What you add The question it must answer Do it yourself when
1 One row per tool call, in your own store What happened, and what did it cost? Always — this is data you cannot buy retroactively
2 An evaluation set of cases with a baseline Did this change help? Under ~50 cases, a script and a CSV are enough
3 Limits and credentials in one place (the gateway) Who may spend what, and with which key? Never — enforcement has to be outside the prompt
4 An orchestrator with durable state and a stop condition Who runs next, and what happens on a crash? For a fixed sequence, a queue and a worker will do
5 Memory, split by kind, with retention and review What is the system allowed to remember? Only the semantic layer needs a product

The ordering is not a preference, it is a dependency: evaluation reads the record, the gateway enforces what the record measures, orchestration writes to the record, and memory is retrieved into it. A team that buys step 4 first usually ends up writing step 1 by hand inside the orchestrator, and paying for the same data twice.

Where SmartGate fits

SmartGate is the place in this stack where the record and the enforcement point are the same object. Every tool call through the gateway is written as an audit row — caller, route, transport, tool, token count, latency, outcome — and the same row is what the per-key rate limit and the per-team token budget read. That is the layer we recommend owning early, and it is the one we operate.

The plan table determines the operational limits rather than the features: monthly token caps are 2M, 20M, 100M and 200M+, requests per minute per key are 120, 300, 600 and 1200, audit-log retention is 7, 30, 90 or 180 days, and a team can hold 2, 10, 30 or unlimited keys. If your retention requirement is driven by an audit or a compliance deadline, compare those numbers against the requirement before you commit a date to them — the pricing page is the authoritative table, and the FinOps lead's view of the same rows is the one that has to survive a budget review.

Frequently Asked Questions

Limitations

This page is a selection guide, not a benchmark: no tool is ranked, because the ranking depends on where your state, your record and your credentials already live. It describes the four layers and the questions that separate the options inside them; it does not claim that any particular product or open-source project implements a layer well.

The plan figures quoted above were re-verified against the live pricing page on 2026-09-30 and are operational limits, not a feature comparison — caps, per-key rates, retention and key counts change with the plan, and a compliance decision should read the current table rather than this page.

This page carries no code excerpt, and the reason is recorded in the Method note below: the slice matcher found no unique symbol for any of its eight sections. Consequences that follow honestly from that: no line-numbered claims, no quoted implementation detail, and no assertion about the internal shape of any framework named above.

Sources

  • Anthropic's engineering note on workflow and agent patterns — Building effective agents, for the workflow-versus-agent split, the supervisor shape and the handoff shape.
  • The Model Context Protocol specification — modelcontextprotocol.io/specification, for how a runtime tool is described and returned, together with the protocol's own introduction at modelcontextprotocol.io/introduction.
  • OpenTelemetry semantic conventions for generative AI systems — opentelemetry.io/docs/specs/semconv/gen-ai, for the span attributes a tracing layer can consume instead of bespoke glue.
  • Langfuse and LangSmith documentation — langfuse.com/docs · docs.smith.langchain.com, as the shape of a tracing and evaluation product in this layer.
  • Temporal's workflow documentation — docs.temporal.io, for durable execution and idempotent steps, which is the property the orchestration section asks a framework to provide.
  • Demand figures in this page are our own measurements: DataForSEO Google Ads, United States, 12-month window, measured 2026-09-30, recorded in this project's search_volume.json and research_brief.md.
  • Product behaviour and the plan table: read from the product source at the revision pinned in this project's pipeline_results.json, read-only, with the plan figures re-verified against the live /pricing page on 2026-09-30.

Method note

This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of 8 sections for this page (0 abstention(s), 8 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the eight section keywords, because this lane's vocabulary — orchestration, tools, evaluation, memory — collides with generic helper and type names across a codebase. A pinned generic name would have given the page the shape of a verified article with none of the substance, so every section above is written from sources.

Product claims were read from the product source at the revision the slice run recorded in this project's pipeline_results.json, read-only, and the plan figures were re-verified against the live pricing page on 2026-09-30. The section keyword quoted above each heading comes from this project's own paid measurement run, not from a third-party tool. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.