SmartGateSmartGate

AI Agent Observability: Steps, Cost, and Failure Traces

AI agent observability is the practice of reconstructing one agent run — not one model call — from the record it leaves behind, so you can answer three questions: what happened step by step, where the tokens and money went, and which step failed.

Short answer: AI agent observability is the practice of reconstructing one agent run — not one model call — from the record it leaves behind, so you can answer three questions: what happened step by step, where the tokens and money went, and which step failed. A run is a sequence of decisions, tool calls and retries, so the unit of analysis is the run and the load-bearing fields are the step, the tool, the tokens, the latency and the outcome.

Key takeaways

  • The run, not the call, is the unit. One user request becomes a plan, several tool calls and often a retry; an observability view built on single calls cannot say whether the run succeeded.
  • Three questions, one record. "What happened", "what did it cost" and "which step failed" are three queries over the same per-step rows, joined by one correlation id.
  • Cost attribution is a grouping choice. The same token count can be read per step, per tool or per session; decide the grain before the invoice forces the question.
  • Failure is a per-step property with a run-level verdict. A retry that succeeds hides a failure at the call level and must stay visible at the run level.
  • Trace data is sensitive by default. Tool arguments and returns carry the payload, so the capture, redaction, retention and access rules are part of the design, not an afterthought.
  • Start with the correlation id. Give every run a trace id when it starts; without it the run cannot be reassembled from rows afterwards.

ai agent observability: three questions, one record

An agent run does not look like a chat completion, and observability built for chat completions does not fit it. A single user request — "summarise these six pages and file the result" — becomes a plan, a search, two fetches, a deduplication pass, a compression pass and a memory write. Six or more calls cross several services, and at most one of them is a model call. The thing a developer actually needs to reason about is the run: the ordered set of steps that turned one intent into one outcome.

Three questions cover almost everything an operator asks about a run, and one record answers all three:

  1. What happened? The step sequence — which tool ran, with which arguments (or a summary of them), in what order, and whether any step was retried or skipped.
  2. Where did the money go? The tokens and the cost, attributed to the step, the tool and the session that produced them, rather than to a monthly total.
  3. Which step failed? The step that errored, the step that was slow, and the step that never ran because an earlier one did.

The three are not three systems; they are three projections of one set of per-step rows joined by a correlation id. The standards body for distributed tracing already models this shape. The OpenTelemetry semantic conventions for generative AI define spans for an agent, for a model inference and for a tool execution, with an explicit agent-versus-tool split, so a run is one trace whose children are steps (GenAI agent spans). The propagation primitive underneath it is W3C Trace Context, which carries a trace id and a span id across process and service boundaries so the pieces can be stitched together afterwards (W3C Trace Context). What those two give you is the skeleton; the fields you attach to each step are what turn the skeleton into answers.

Question The field it reads The query it becomes
What happened? step index, tool name, argument summary, status the ordered step list for one trace id
Where did the money go? input tokens, output tokens, cost, model a sum grouped by step, by tool, or by session
Which step failed? status, error class, duration per step the first non-success span, plus the step that never started
How slow was it? per-step duration, wall-clock per run duration distribution and the p95 tail per step

The record shape for a single call, field by field, is the subject of LLM observability: what one tool call should leave behind — this page takes that row as its input and reads it as a sequence. The distinction matters because a row is a fact and a run is an argument: only the sequence tells you whether the agent chose the right tool, whether a retry masked a failure, or whether the run stopped because it was done or because it ran out of budget.

llmops: the agent run as the unit of analysis

LLMOps is the practice of running language-model systems as production software, and its centre of gravity shifts when the system is an agent. In a chat product, one request maps to one model call, so the request is both the reliability unit and the cost unit. In an agent, both units move up a level. Reliability is a property of the run: a run that took nine steps instead of three is not broken, but it is slower, more expensive and more likely to have hit a transient error along the Cost is a property of the run too, because retries and re-reads multiply the token count in a way a single call cannot show; token optimization techniques address the number itself.

That shift is why "LLMOps" and "observability" are often used as if they were the same word, and why the agent layer has to be separated from the model layer. The model layer answers questions about a completion: how many tokens, how much latency, which model, what the output looked like. The agent layer answers questions about a run: which steps ran, in what order, at what cost, and where it stopped. The two are complements, and the ordering is deliberate — instrument the run first, because the cost, reliability and incident questions all resolve against the run, and the per-completion detail is a zoom-in on one of its steps.

Two design decisions make the run analysable, and both are cheap at design time and expensive to retrofit. The first is a correlation id: mint a trace id when the run starts, propagate it into every step, and store it on every row. Without it, "the run" is something you guess from timestamps and adjacent rows, which fails exactly when concurrency makes rows interleave — the reason concurrency control in AI backend systems is a precondition for the trace, not an optimisation on top of it. The second is a step index: record the position of each step in the run, so an out-of-order arrival at the store can still be reassembled into a sequence, and so "the third step" is a stable identity in a review.

A run also needs an explicit stop condition that is not the model declaring itself finished. A maximum step count, a token budget for the run, and a wall-clock ceiling are the three that cover most runaway cases; each is a field the run record can carry, which is what turns "the agent looped for an hour" from a surprise into an alert.

agent evaluation: read the trace before you trust the answer

Agent evaluation and agent observability are usually discussed as separate disciplines, and in practice they share one artifact: the trace. You cannot judge whether an agent is getting better without knowing what it did, and the steps are where the difference between a lucky answer and a reliable one lives. A run can return a correct final answer after choosing the wrong tool, retrying three times and burning four times the tokens — a single output score cannot see any of that.

The evaluation method that has converged in the field is case-based and run-level. Anthropic's builder guidance on agent evals describes evaluation sets of cases with checkable properties, run against the whole pipeline rather than the model in isolation, precisely because most regressions arrive from a changed tool description, a new retrieval result or a truncated context rather than from the model itself (Demystifying evals for AI agents). The trace is what makes that practical: each case produces a run, and the run produces the numbers an evaluation reads — steps taken, tools chosen, retries, tokens, and the outcome of each step.

Three trace-derived signals carry most of the value, and none of them is the final answer text. Step count against a baseline catches the agent that got slower or loopier after a prompt change. Tool selection catches the agent that used search where it should have used a fetch — a defect invisible in the output. And retry structure separates a run that succeeded first try from a run that succeeded after two failures; the second is a latent bug wearing a green checkmark. LangSmith's evaluation documentation frames the same split between judging an output and judging a trajectory, which is the run-level view (LangSmith evaluation).

The honest boundary is worth stating: observability tells you what happened; evaluation tells you whether what happened was good. The trace is the input to evaluation, not a substitute for it, and a team that instruments the run without defining cases has a rich record and no way to tell a regression from a rewrite. The cluster centre page covers the record itself; the question of which product reads it best is how to evaluate and shortlist LLM observability tools.

Where the run id comes from, read from our own implementation

A run view needs an identifier that both sides agree on. Here is how ours is produced, read on 2026-10-07 from backend/smartgate/core/mcp_session_trace.py and backend/smartgate/core/events.py.

  • One session, one stable trace. The gateway keeps a session-to-trace map: an MCP session that arrives without a trace gets an identifier of its own, and every call in that session reuses it, so a run can be reassembled from the calls rather than reconstructed from timestamps.
  • A client-supplied trace wins, and is remembered. If the calling agent sends its own trace header, that value is used and stored against the session. This is the difference between a trace id that exists only in our logs and one that appears in both systems, which is what makes cross-tool correlation possible without a join on time.
  • The honest limit: the map is per process. It lives in memory, so a worker restart loses the association — an id that must survive a restart has to come from the caller's header. Plan for that before you build a dashboard on top of it, or you will debug a correlation gap that is not a bug.
  • In-flight signals are events, not rows. A small publish/subscribe bus with an async queue and a single worker carries internal signals: subscribers register by event name, callbacks may be sync or async, and an exception inside one subscriber is caught so a single bad consumer cannot stall the stream. The durable per-call record is the audit row; the bus is the fast, lossy layer above it — and a run view is assembled from both.

ai agent monitoring: latency budgets and failure modes

Monitoring an agent means watching the run's shape over time, not just its error rate, because an agent degrades before it fails. The signals that move first are structural: steps per run creeping up, retries per run rising, time-to-first-tool growing, and the tool mix shifting. Each is a query over the step rows, and each catches a different class of problem before it reaches a user.

Latency deserves its own treatment because a run's latency is not one number. It is the sum of the per-step durations plus the waiting between them, and the useful cuts are at least three: the model inference time per step, the tool execution time per step, and the queueing or dispatch time between steps. A run that is slow because one tool is slow is a different problem from a run that is slow because the planner emitted nine steps, and the two look identical if you only measure end-to-end. OpenTelemetry's generative-AI metrics define duration and token-usage instruments for model, agent and tool operations, which is the vocabulary that makes the breakdown a dashboard rather than a bespoke script (GenAI metrics). The tail is where the incident lives: a p50 run and a p99 run of the same agent often differ by a retry storm, and the p99 is the number the on-call engineer is paged for.

Failure modes cluster into five classes, and naming them makes them countable:

  • Tool error — the call reached the tool and the tool failed; the run may retry or abort.
  • Model error — a timeout, a rate limit or a malformed response from the inference step.
  • Planner no-progress — the agent repeats a step or oscillates between two, never advancing.
  • Budget exhaustion — the run hits its token or step ceiling and stops early, often with partial work that was still paid for.
  • Partial failure — some steps succeed and one fails, leaving a result that looks complete.

The pattern that makes these visible is to record the outcome at the step level and derive the run verdict from the steps, rather than overwriting a run with a single status. A run with two successful calls and no third is a different event from a run whose second call errored and was retried — both are one query away if the steps are preserved, and neither is visible if the run is flattened to a boolean. The event vocabulary for signalling these transitions as they happen comes from the same generative-AI conventions (GenAI events), and the choice of what to sample versus keep is the one the tool-selection page takes up. Where the record is written cheaply enough to keep at full volume is a property of the low-cost AI backend architecture the agent runs on.

langfuse alternative: what an agent-run view adds to prompt tracing

If you already run prompt-level tracing, the honest question is not which dashboard is prettier but which events exist in the record. The two approaches overlap on the model call — both can show a completion's tokens and latency — and diverge on everything a run is made of. Prompt tracing is centred on the generation; an agent-run view is centred on the step tree, and the step tree is what carries the questions "which tool did it choose" and "why did it stop".

The practical consequence is a difference in what can be reconstructed. A prompt-level trace can show you a model call's tokens; it usually cannot show you that a search and two fetches happened at all, that one of them was retried, or that the third step never ran. An agent-run view stores the tool executions as first-class steps, so the run is complete even when most of its steps are not model calls. That is the same split the OpenTelemetry conventions formalise when they give an agent span tool-execution children rather than folding tools into the model span.

The other property of the agent layer is that it can be vendor-neutral. Modern agent stacks emit OpenTelemetry-compatible traces, and open instrumentation projects — OpenLLMetry for the model and framework calls, OpenInference for the span conventions an evaluation tool reads — let the same trace feed an observability product, a local viewer or a data warehouse without rewriting the agent (OpenLLMetry, OpenInference). The reason to care at the agent layer specifically is that a run is long-lived and multi-service: the trace outlives any single process, so an open format is what keeps the record readable when the framework or the vendor changes. Which of those products to standardise on is a selection decision rather than an agent-layer one, and it belongs to the sibling page above.

langsmith alternative: trace data and the privacy boundary

A trace is more sensitive than a log line, and that is the fact most observability rollouts discover late. To reconstruct what happened, a useful trace captures tool arguments and, often, tool return values — which is to say it captures the data the agent was handling. A summarisation run's trace can contain the document; a support run's trace can contain the customer's message; a database run's trace can contain the row. Treating that record as ordinary telemetry is how a debugging tool becomes a data-residency problem.

The boundary has four decisions, and they are cheaper to make before the first run than after the first incident:

  • Capture. Decide per field whether the trace stores the value, a hash, a redacted form, or only the metadata — the tool name, the token count, the duration and the status. Metadata-first capture keeps the run reconstructable while leaving payloads out of the default record.
  • Redaction. Where a payload must be kept, strip identifiers at the point of capture rather than at the point of query, so the sensitive value never lands in the store at all.
  • Retention. Set a window per class of data and delete on schedule; a trace store that keeps everything forever is a liability with a query interface. Plan-level retention in this stack is a configured number of days rather than a fixed constant, so the current pricing page is the authority on what a given tier keeps.
  • Access. Scope every read to the tenant that owns the run. A trace id that works as a bearer token across tenants is an access-control hole; tenancy has to be a mandatory filter on the query, not a field in the payload.

Two external pressures make this concrete rather than theoretical. The EU AI Act's record-keeping provisions require automatic logging of events over the lifetime of certain high-risk systems, which sets a floor on what must be captured even as privacy sets a ceiling on what may be (EU AI Act Article 12). And the NIST AI Risk Management Framework asks for continuous monitoring and traceability, which is a standing demand for the record rather than a one-off audit (NIST AI RMF). The two pull in the same direction as the engineering advice: keep the metadata that makes the run auditable, and treat the payload under an explicit policy. When one gateway serves many customer workspaces, the boundary is also a multi-tenant one — the pattern is the one described for AI automation for SaaS operations.

A related boundary is trust. An agent's tool returns are attacker-influenced input as far as the model is concerned, and the Model Context Protocol's security guidance is explicit that servers and clients must not treat tool output as trusted (MCP security best practices). For observability that has a direct consequence: the trace is a record of untrusted content, so it must not be rendered or re-fed without the same care the live data gets. Capturing a tool return into a trace and then displaying it in a shared dashboard moves untrusted content into a new sink.

How to get started

The first move is not a product; it is a correlation id and one row per step. Everything else reads from that.

  1. Mint a trace id per run and propagate it. At the start of the run, create an id and pass it into every step; store it on every row. This one field is what makes the run reassemblable.
  2. Write one row per step, with a step index. Tool name, argument summary, input and output tokens, duration, status and the trace id. Record the attempt even when the step failed, because the failed attempt is the interesting one.
  3. Record tokens and cost where the step completes. Keep the count on failed steps too — work that was paid for and discarded is exactly what a cost review looks for.
  4. Set a stop condition per run. A maximum step count and a token ceiling, carried as fields, so a runaway loop degrades to a bounded event instead of an invoice.
  5. Decide the privacy boundary before volume arrives. Metadata by default, payloads on exception, a retention window per class, and a tenant filter on every query.
  6. Add a run-level view over the rows. The step sequence, the cost grouped by tool, and the failed-step count are three queries, and they are the three answers the on-call engineer needs.

Where SmartGate fits

SmartGate is the control plane an agent run passes through, made concrete: an MCP-native algorithm gateway for token control, traffic shaping and agent audit. It is a single authenticated endpoint through which an agent reaches its tools, with per-key metering and an audit row written as calls happen. Seven tools are exposed through it — smart_fetch, smart_search, smart_context_gate, smart_dedup, smart_budget_guard, smart_memory and smart_pipe — and each call is counted against the caller's key.

It is worth being precise about what this contributes to agent observability and what it does not. SmartGate does not inspect model internals and does not claim to store or reconstruct prompts; it records the call — the tool, the token count, the latency and the outcome — under the key that made it. That is the agent-layer record described above for every call that routes through the gateway, and it is what makes the per-tool and per-session cost cuts possible without a second metering system. Traffic shaping bounds what a looping or runaway agent can spend while it is spending it, which is the control-plane answer to the same failure modes monitoring observes after the fact.

The plan table determines the operational limits rather than the feature set: monthly token caps of 2M, 20M, 100M and 200M+, requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and a team holding 2, 10, 30 or 9999 keys. Read those against your own retention and attribution requirements on the pricing page, which is the authoritative table. The call control lives on token control and the record on audit and compliance; to watch one run end to end, start free and follow a single trace from the first tool call to the audit row it wrote.

Frequently Asked Questions

What is the single most important field for agent observability?

The correlation id, usually called a trace id. One agent task is many steps spread across services, and the trace id is the only field that reassembles them into one run. Mint it when the run starts, propagate it into every step, and store it on every row; without it the run has to be guessed from timestamps, which fails as soon as two runs overlap.

Should I measure cost per call or per run?

Both, but the run is the decision unit. A per-call cost tells you which tool is expensive; a per-run cost tells you what the task actually cost, including retries and re-reads, and that is the number a budget or a pricing decision reads. Store the per-step tokens and derive both aggregations from the same rows rather than metering the two separately.

How is a failed retry different from a failure?

A retry that succeeds leaves a green run with a hidden failure inside it. That is why the step rows must be preserved: a run with one error span followed by a success is a latent reliability problem, while a run with a single success and no error is healthy. Collapsing the run to a final status deletes exactly the signal an incident review needs.

Do I need to store prompts and tool returns to debug a run?

Usually not. Metadata — tool, tokens, duration, status, trace id — reconstructs what happened, where the cost went and which step failed without storing the payload. Store payloads only where a specific case requires them, with redaction, a retention window and a tenant-scoped read. A trace store that keeps everything is a data-residency exposure with a query interface.

Does an observability tool stop a runaway agent?

Not by itself. Observability shows the loop after the fact; the stop happens at a control point that can refuse or throttle a call while it is running. The durable design pairs the record with a gateway that enforces per-key limits and a run-level budget ceiling, so the failure mode is bounded at the point it would otherwise be discovered on an invoice.

Limitations

This page is a design frame for the agent layer, and it is deliberately bounded. It assumes a per-step record already exists or can be built; it does not benchmark any observability product, and it makes no claim about which one reads a trace best — that is a selection exercise with its own criteria and its own page. Where a number appears it is either a plan limit this project records or a property the cited source publishes; no latency, cost or reliability figure is invented for any agent.

The advice has real limits worth naming. A trace shows what an agent did and not why the model chose it, so the trace does not replace an explanation of a bad decision. Metadata-first capture means the default record cannot answer "what exactly did the model see", which is a payload question with different retention and privacy constraints. And full-volume, one-row-per-step records are cheap per row but additive at scale, so storage has to be budgeted as deliberately as tokens are. Where a regulation or a contract sets a retention or residency requirement, that requirement decides the policy rather than this page.

Sources

Method note

This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned none of this page's seven sections (0 abstentions, 7 no-slice verdicts): rule A found no unique symbol in the scanned repository for any section keyword, and the remote candidate fallback returned generic helpers that are collisions rather than section-specific evidence. A pinned generic would have given the page the shape of a verified article with none of the substance, so every section above but one is written from external, linkable sources, which is the house rule for an unpinned section. The exception is "Where the run id comes from, read from our own implementation" — our own trace handling, read on 2026-10-07 from backend/smartgate/core/mcp_session_trace.py and backend/smartgate/core/events.py: the session to trace mapping, the client-header precedence, the per-process limit, and the event bus.

Product facts were read read-only from the product source at the revision recorded in this project's pipeline_results.json, and the plan figures were re-checked against the live pricing page on 2026-10-04; the demand figures are this project's own measurement. Every external statement quoted above is taken from the URL cited beside it, and each external URL returned HTTP 200 when it was fetched before publication. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.

The slice run for this page recorded 0 of 7 sections pinned, 0 abstention(s) and 7 no-slice verdict(s); BLOCKS is empty because the matcher found no unique symbol for any section rather than section-specific evidence, as the Method note above explains.