LLM Observability Tools: How to Evaluate and Shortlist
Most LLM observability tools collect spans, draw dashboards and alert on the same events, so a feature grid will not separate them. Four axes will: the sampling policy and its error tail, the cost model, the lookback window, and whether the tool can feed a per-team quota while the call is still running.
Short answer: Most LLM observability tools collect spans, draw dashboards and alert on the same events, so a feature grid will not separate them. Four axes will: the sampling policy and its error tail, the cost model, the lookback window, and whether the tool can feed a per-team quota while the call is still running.
Key takeaways
- Feature lists converge; contracts do not. Sampling, pricing, lookback and the quota handoff are the four properties a buyer cannot change after signing, so they are the four worth testing.
- A sample you cannot reweight is a guess. Ask what fraction of spans lands and whether errors are exempt; an unsampled error count and a sampled one are different numbers on purpose.
- Storage is the second invoice. Per-span pricing hides the cost of the bytes you keep — prompt and completion text is usually the line nobody budgeted.
- Lookback is an operational ceiling, not a storage footnote. It sets how long after an incident you can still answer a question with data rather than memory.
- The quota handoff is the dividing line. A tool that only aggregates after the fact observes spend; one that returns a number mid-call can help control it.
- Do this first: take the record you already produce, run two weeks of it through the candidate's free tier, and check whether the four answers above hold before you compare a single screenshot.
llm observability tools: the four axes, not the feature list
Every serious candidate in this category will show you the same three capabilities: it collects spans from your agents, it renders traces and dashboards, and it alerts when an error rate moves. Those are table stakes, and a comparison built on them produces a tie. The axes that actually separate candidates are the ones that decide what the tool costs you in money, in fidelity and in operational freedom:
| Axis | The question that discriminates | Why it decides the purchase |
|---|---|---|
| Sampling | What fraction of spans lands, and are errors exempt from the sample? | A sample you cannot reweight turns every cost report into an estimate |
| Cost model | Per span, per host, per seat, or per stored byte? | The second invoice — payload storage — is the one nobody budgets |
| Lookback | How far back will the vendor's own query actually reach? | It sets how long you can answer an incident with data, not memory |
| Quota handoff | Can the tool return a number before the call completes? | It is the line between observing spend and controlling it |
The reason to lead with these four is that they are the ones a reader cannot change after signing. A missing chart is a feature request; a sampling policy that discards the exact traces you need, or a fourteen-day lookback when your incident review runs monthly, is a contract decision. Everything else in a vendor deck — the heat map, the prompt playground, the evaluation runner — is a feature of one of the four axes, and it is cheaper to add later than to migrate for.
There is also a reason the axe falls here rather than on price alone. Two tools can carry the same list price and differ by an order of magnitude in what they cost to operate, because the expensive part of observability is rarely the collector: it is the volume of spans you choose to retain and the size of the payloads attached to them. A buyer who tests those two numbers tests the bill.
llm observability: what a candidate must be able to read
The first filter is not a dashboard at all. It is whether the candidate can read the record your system already produces. If one row exists per tool call — the tenant or team it belongs to, the tool name, the tokens it consumed, its latency and its outcome — then any serious product can ingest it, and the comparison starts from data you own. If your agent framework emits OpenTelemetry spans, the same test applies one level up: a product that consumes the generative-AI span conventions without bespoke glue saves you an integration project before the pilot starts. The shape of that record, field by field, is what LLM observability sets out; this page takes it as the input contract and asks which tool reads it best.
Three read requirements follow from that contract, and each is checkable in an afternoon.
- It must accept the fields you have, not only the fields it wants. A tool that insists on its own SDK in the hot path adds a dependency to every call; a tool that accepts a span over the wire or a log line keeps the collector optional.
- It must keep the tenant on the row. Team or workspace is the dimension every cost, access and abuse question is asked against; a product that wants to reconstruct tenancy by joining to your users table has moved work back into your codebase.
- It must be able to read a correlation key. One agent task is several calls, and the tool that can group them answers "what did this task cost" where a tool without it can only answer "what did this call cost".
A candidate that passes the read test but fails any of the four axes is still a candidate; one that fails the read test is not, because the migration cost lands on your team rather than the vendor's.
agent observability: sampling rate and the error tail
Sampling is the first axis because it is the one that quietly changes every other number. At a high volume of spans, no vendor keeps all of them, and the honest question is not "do you sample" — the answer is almost always yes — but which spans survive. Three policies are common and they are not interchangeable. Head-based sampling decides at the start of a trace and is cheap, but it is blind: it keeps or drops a whole task before anyone knows whether it failed. Tail-based sampling decides after the trace ends and can therefore keep every error and every slow outlier while dropping the boring majority, which is what makes the sampled data usable for an incident. Hybrid policies sample aggressively at the head and then force-keep anything that errored, which is usually the right default for agent work.
The test to run is a pair of numbers, not a policy name. Ask the candidate what fraction of spans it retains at your volume, and separately how many errors it keeps. A tool that answers "we sample at ten percent" has told you that your error rate is measured on a tenth of your failures, and if it cannot reweight the sample back to a population estimate, then every chart downstream is an estimate wearing a decimal point. The reverse is also true: a tool that keeps all errors and samples the happy path gives you a free way to compute an error rate — count the kept errors, multiply the successes by the sampling factor.
One more property belongs to this axis. Sampling interacts with load, and agent traffic arrives interleaved rather than in task order, which is why concurrency control in AI backend systems is a precondition for the correlation key a sampling policy depends on. If two workers interleave their spans, tail-based sampling has to buffer the whole task before it can decide, and the buffer is a resource the buyer ends up paying for somewhere.
ai agent monitoring: the cost model, and who pays
The second axis is what the vendor charges for, and the useful discipline is to price the tool on your own numbers rather than the vendor's tier names. Four pricing shapes dominate: per span ingested, per host or per agent instrumented, per seat, and per byte stored. Each one pushes a different behaviour onto the buyer.
| Pricing shape | It rewards | It punishes | Ask before signing |
|---|---|---|---|
| Per span | trimming what you send | verbose agents, many tool calls per task | the included span count and the overage rate |
| Per host | fewer, bigger collectors | scale-out and many small services | whether autoscaling hosts are billed |
| Per seat | many spans per engineer | a small team on a large estate | who counts as a seat |
| Per byte stored | metadata-only retention | prompt and completion payloads | the retention tier that triggers it |
The trap is that the visible price is almost never the whole invoice. A per-span plan is cheap until an agent that used to take three calls takes nine, because span count tracks agent verbosity rather than user count; a per-seat plan is cheap until a platform team of four owns an estate of sixty services. The predictable hidden line is storage: a tool that retains prompt and completion text at per-byte pricing is selling you a data-residency commitment with a monthly meter, and the two levers that shrink it are the same ones that shrink the model bill — fewer bytes on the wire and a cheaper model on the easy steps, which is what token optimization techniques deals in.
Cost control inside the tool is the other half. The question to ask is whether the candidate can turn a measurement into a target: a per-task budget rather than a monthly total, an alert before the budget rather than after it, and a breakdown by tool or step rather than by model alone. That is the work optimize agent execution cost describes, and it is the difference between a tool that reports an alarming number and one that changes it. A buyer comparing candidates on the shape of that breakdown is comparing the two products on the only report someone will act on.
mcp monitoring: lookback window, retention and the export path
The third axis is how far back the tool will reach, and it is the axis most often read off a storage footnote when it is really an operational ceiling. "Thirty days of retention" can mean any of three different things, and they are worth separating before they are compared. Query lookback is how far back the interactive product will actually run a filter — the window an engineer has on a Tuesday afternoon. Export reach is how much history a bulk read can pull for a SIEM or a review, which is often longer because it is slower. Legal retention is how long the vendor will keep the data at all, whether or not the console surfaces it. A candidate can be generous on one and stingy on the others, and the reader who assumes the numbers agree is the one surprised by an incident review.
For MCP-shaped traffic the same three questions apply with one addition: what counts as a unit. A
client opens one stream and calls several tools inside it, so a tool that meters or retains per
session has to decide whether a five-tool conversation is one record or five. The answer the buyer
wants is five, because the failed call is the interesting one, and a product that collapses a session
into a single event loses the exact detail an MCP monitoring view exists to show. This is
also where the product's own plan knobs matter: retention in this stack is a per-plan value —
seven, thirty, ninety or one hundred and eighty days by tier — and the current table is the
authoritative one, so treat any figure quoted second-hand (including on this page) as a reading to
re-check against /pricing.
The check that makes the axis concrete is one query. Take an incident from two months ago, open the candidate's free tier, and ask it to return that window. If the answer is "not on this plan", you have learned the retention that matters, and you have learned it before the invoice rather than after.
Quota integration: can the tool feed a per-team limit?
The fourth axis is the one that decides whether the tool is a reporting layer or a control point, and it turns on a single question: can the candidate hand a number back before the call finishes? An observability product that aggregates spans after the fact tells you what a team spent last month, which is valuable and late. A product that can answer "how many tokens has this team committed this period, and what is its ceiling" while the request is still open is something else — it is the input to a decision that can refuse the call. The mechanics of that decision belong to how to enforce a token quota per team; what a buyer tests here is whether the tool exposes what enforcement needs.
Three interface properties decide it. A read at request time — the counter or entitlement has to be queryable inside the latency budget of the call, not only in a nightly export. A stable attribution key — the team id the tool records has to be the same id the limiter resolves, or the two systems will disagree by exactly the amount nobody can explain. And one definition of the period — a tool that reports a rolling thirty days while the limit resets on the calendar month will look wrong once a month, every month. A candidate that satisfies all three can sit in front of the model and tool calls and be part of the control loop; a candidate that satisfies none of them is a dashboard downstream of enforcement, which is a legitimate purchase but a different one.
It is worth being blunt about which of the two the reader usually wants. If the goal is a report for a review, aggregation is enough. If the goal is to stop a team before it spends the month, the tool has to be readable at the moment of the call, and the shortest route is usually to buy the record and the limiter from the same operator rather than to stitch two products across an API boundary.
Who reads the same rows: the budget review
The audience for a tool's output is not only the engineer on call, and a shortlist that forgets the second reader tends to pick a product the organisation cannot use. The second reader is whoever owns the bill, and they ask a different set of questions: which team or feature drove the change in spend, what is the trend against the plan, and what is the per-unit cost that procurement can hold someone to. Those are aggregations over the same rows the engineer queries, viewed at a coarser grain.
Two design properties decide whether a candidate can serve both readers from one dataset. The rows must carry stable dimensions — team, tool, model, feature — rather than only an opaque trace id, because a budget review groups by exactly those and nothing else. And the tool must allow a different time grain without a different pipeline: a monthly trend and a per-minute error rate should be two queries over one store, not two copies of the data that drift. A product that solves the engineer's problem with unstructured payloads and the finance reader's problem with a separate reporting export has quietly built the two-definitions-of-a-month defect into the contract. The reader who owns spend rather than enforcement will recognise the same figures framed for a review in the FinOps lead's view of the same rows.
A scorecard you can fill in one afternoon
Written as questions a buyer can answer from a trial account and two weeks of real traffic, so the shortlist is grounded rather than persuasive.
| Axis | Question to ask the vendor | The answer that should worry you |
|---|---|---|
| Sampling | What fraction of spans is retained at our volume, and are all errors kept? | "We sample" with no error exemption and no reweighting |
| Sampling | Can the kept sample be scaled back to a population estimate? | The chart shows a raw count with no sampling factor |
| Cost | What does one prompt-plus-completion payload cost to store for a month? | Payload storage is "included" with no volume stated |
| Cost | Which dimension does the bill track — spans, hosts, seats or bytes? | The overage rate is undisclosed until invoicing |
| Lookback | Run a query from two months ago in the trial. | The window is shorter than our incident review cadence |
| Lookback | What does a bulk export reach, and at what cap? | History older than the console is simply gone |
| Quota | Can we read this team's committed tokens at request time? | Only a nightly aggregate is available |
| Quota | Does the team id match the one our limiter resolves? | Attribution is reconstructed from a users table |
| Export | Is the export ascending, capped and complete? | Exports truncate silently at a fixed row count |
The scorecard is deliberately about what the vendor will not change, not about what looks good. A candidate can score poorly on a slide and still be the right purchase if the two axes you actually depend on are strong; the failure mode is choosing on the strongest demo and discovering the weakest contract term in month two.
The shortlist order: five checks before you sign
The order matters because each step makes the next one cheaper to run.
- Take stock of the record you already keep. If one row per tool call exists, you can trial any candidate without an integration project; if it does not, the first purchase is the record, not the dashboard.
- Trial one candidate on real traffic, not a demo tenant. Two weeks of your own spans answers the sampling and cost questions in a way a sales figure never will.
- Ask one incident question from two months back and time the answer. This settles lookback and exposes the export path in a single step.
- Test the quota handoff in the latency budget. Issue a read at request time and confirm it returns inside the call; a number that only exists later is a report, not a control.
- Read the retention and overage terms against your own obligations. Retention is a policy value per plan on this stack, and the outside auditor asks about it in days, not in tiers.
Only after those five do the interface questions — the operator UX, the query language, the alerting — earn much weight, because they are the ones a team can work around. A tool that wins on step five and loses on step three is still the wrong contract.
Where SmartGate fits
SmartGate is the operator in this stack where the record and the enforcement point are the same object, which is what makes it a candidate under the fourth axis above rather than only the first. Every tool call through the gateway is written as one audit row — caller, route, transport, tool, token count, latency and outcome — and the same row is what a per-key limit and a per-team token budget read. For a buyer, that means the quota handoff is not an integration between two products but a read of the row that already exists; for the reader comparing candidates, it is the shortest answer to the scorecard's quota rows.
The plan table is the part that decides the operational limits rather than the features, and it is
worth reading as a contract, with the current pricing page as the authority. Monthly token caps are
2M, 20M, 100M and 200M+; requests per minute per key are 120, 300, 600 and 1200; audit-log retention
is 7, 30, 90 or 180 days; and a team can hold 2, 10, 30 or 9999 keys. Quote those as the shape of the
current table and re-check the numbers on /pricing before committing a deadline to them, because
retention in days is exactly the term a compliance review reads first. The same figures appear as the
spend story a finance reader wants in the FinOps lead material, and as the underlying rows an
engineer queries in LLM observability.
Frequently Asked Questions
Limitations
This page is an evaluation frame, not a benchmark: no product is ranked, because the right candidate depends on where your record, your credentials and your limits already live. It describes the four axes and the questions that separate candidates inside them; it does not claim that any named product implements an axis well, and it does not enumerate features.
The plan figures quoted here were re-verified against the live pricing page on 2026-09-30 and are operational limits, not a feature comparison — caps, per-key rates, retention and key counts move with the plan, so a compliance or procurement decision should read the current table rather than this page.
This page carries no code excerpt, and the reason is recorded in the Method note below: the slice matcher found no unique symbol for any of its sections. The honest consequences are that there are no line-numbered claims and no quoted implementation detail about any product named above. Cost figures are absent on purpose: pricing shapes are described structurally, and the reader is expected to price a candidate on their own traffic rather than on a number reprinted here.
Sources
- OpenTelemetry, trace sampling concepts (head-based and tail-based sampling) — opentelemetry.io/docs/concepts/sampling, for the sampling policies the first axis compares and the reweighting problem they create.
- OpenTelemetry semantic conventions for generative AI systems — opentelemetry.io/docs/specs/semconv/gen-ai, for the span attributes a candidate can read without bespoke glue.
- Model Context Protocol specification — modelcontextprotocol.io/specification, for the stream-and-tool-call shape that decides whether a five-tool session is one record or five.
- Langfuse documentation — langfuse.com/docs, as the shape of a tracing and evaluation product whose pricing is metered by ingested events.
- Datadog AI and agent observability — docs.datadoghq.com, as the shape of a host-and-span-metered product in the same market.
- Honeycomb, sampling and event-volume pricing — docs.honeycomb.io, for how a per-event model addresses the sampling and cost axes together.
- Demand figures on this page are our own measurements: DataForSEO Google Ads, United States,
location 2840, measured 2026-09-30, recorded in this project's
search_volume.jsonandresearch_brief.md. - Product behaviour and the plan table: read from the product source at the revision pinned in this
project's
pipeline_results.json, read-only, with the plan figures re-verified against the live/pricingpage on 2026-09-30.
Method note
This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher pinned 0 of 5 sections for this page (0 abstention(s), 5 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the five section keywords, because this lane's vocabulary — observability, monitoring, tools — collides with generic recorder, logger and helper names across a codebase. A pinned generic name would have given the page the shape of a verified article with none of the substance, so every section above is written from sources.
Product claims were read from the product source at the revision the slice run recorded in this
project's pipeline_results.json, read-only, and the plan figures were re-verified against the live
pricing page on 2026-09-30. The section keyword behind each heading comes from this project's own paid
measurement run, not from a third-party tool. No code, batch fingerprints, auction data or internal
hosts are transcribed, so there is nothing here that has to be asserted verbatim.