SmartGate

LLM Observability: What to Record on Every Tool Call

In an agent system the observability unit is the tool call, not the model turn. A record worth keeping has eight fields — a timestamp, the team that owns the request, a request id, the tool name, the arguments, the tokens that call consumed, its latency and whether it succeeded — and everything else (a task view, a trace view, a cost report, an export to your SIEM) is a query over those eight.

Short answer: In an agent system the observability unit is the tool call, not the model turn. A record worth keeping has eight fields — a timestamp, the team that owns the request, a request id, the tool name, the arguments, the tokens that call consumed, its latency and whether it succeeded — and everything else (a task view, a trace view, a cost report, an export to your SIEM) is a query over those eight. Systems that only trace prompts and completions can show you what the model said; they cannot show you which tool call spent the money or which one your security team needs to look at.

Key takeaways

  • Log the tool call, not the conversation. The call is the unit that has a name, a cost, a duration and an outcome; the conversation has none of those properties.
  • Eight fields are enough to answer four different questions. The same row serves monitoring (how many calls failed), governance (what happened between these two timestamps), cost (which tool is expensive) and security (who called what, from where).
  • A trace id turns N rows into one task. Multi-step agent work is one user request spread across several calls; without a correlation key the task view has to be invented later.
  • Tenant isolation is a query property, not a dashboard setting. If the team filter is not part of every query's WHERE clause, an observability feature is a cross-tenant leak.
  • Start by writing down the eight fields for one tool call today — a log line, a table, a JSON blob, anything — before you shop for a dashboard to render them.

Who this page is for

This is written for the engineer who has agent logs spread across two tool calls and a container stdout, and for the platform owner who has to answer "what did the agent actually do at 14:20?" without reading source code. It is not a tour of tracing vendors: it describes the record, the queries you will want against it, and the failure modes that only appear once more than one team or more than one hop is involved. The examples come from a gateway that runs agent tool calls in production, so the field names are real and the queries are the ones the product actually issues.

Why prompt-level tracing is only half the picture

The first generation of LLM observability was built around the generation call: prompt in, completion out, tokens counted, latency measured, cost estimated. That is a good fit for a chat product, where one request maps to one model call. It maps poorly onto agents, where one request becomes a search, two fetches, a deduplication pass, a compression pass and a memory write — six calls across four services, of which at most one is a model call.

Three external developments push the same way. The OpenTelemetry GenAI semantic conventions extend the tracing data model with execute_tool spans alongside model spans, which is the standards body saying that a tool invocation is a first-class span. The NIST AI Risk Management Framework asks for continuous monitoring and incident traceability rather than a periodic review — the evidence has to exist at the moment of the event. And the EU AI Act logging obligations for high-risk systems, applicable from August 2026, make automatic recording of events a legal requirement in that scope, not a debugging convenience. None of those tell you which fields to store; that part is decided by the queries you need.

LLM observability starts with the record, not the dashboard

A dashboard is a rendering of rows. If the rows are right, several dashboards become possible; if the rows are wrong, no dashboard can fix it. The gateway's own audit logger is a good worked example because it is deliberately boring — one table, no event bus, no sampling. It creates the table on first use:

# backend/smartgate/core/audit.py — source lines 39–49 (AuditLogger)
            CREATE TABLE IF NOT EXISTS audit_logs (
                id SERIAL PRIMARY KEY,
                timestamp TIMESTAMPTZ NOT NULL DEFAULT NOW(),
                team_id VARCHAR(64) NOT NULL,
                request_id VARCHAR(64) NOT NULL,
                tool VARCHAR(32) NOT NULL,
                params JSONB DEFAULT '{}',
                token_used INT DEFAULT 0,
                latency_ms FLOAT DEFAULT 0.0,
                success BOOLEAN DEFAULT TRUE
            )

Read that schema as a contract. team_id makes the row attributable to a tenant and therefore queryable without a join to a users table. request_id links the row to the HTTP request that caused it. tool is a bounded string — one of seven algorithm primitives rather than free text, which is what makes "which tool is slow this week" a GROUP BY and not an NLP problem. params is JSONB, so the arguments stay inspectable without a schema migration per tool. token_used and latency_ms carry the two numbers every cost and SLO conversation needs. success is the boolean that turns a log into an error rate. The index on (team_id, timestamp DESC) is the shape of the read path: "this team, most recent first".

Two design notes that are easy to get wrong. First, the record is written by the gateway, not by the tool: a tool cannot forget to instrument itself, and a new tool inherits the fields for free. Second, there is no prompt or completion text in the row. The audit trail answers what was called, by whom, at what cost; storing the payloads as well would turn a compliance table into a data-residency problem, and the arguments are frequently the sensitive part. Endpoint-level traces cover the payload question separately, with their own retention and access rules.

AI agent tools: the audit tag has to survive the transport

The awkward case is the tool that does not know it is being audited. With the Model Context Protocol, a client such as an IDE extension opens one stream and calls several tools inside it; the arguments arrive mid-stream, and the transport context that carries the tenant is established once, before the stream begins. A record written at that point can easily lose the tag it needs.

The mechanism that keeps it honest is a re-bind: before each in-stream call executes, the audit context is re-attached from the ambient request context rather than threaded through every call site. The practical consequences are worth stating plainly, because they decide how the logs read:

  • Every call in the stream is its own row. A five-tool conversation is five rows sharing one request_id, not one row with an array of calls.
  • A transport failure still leaves a trail. The tool may not have run, but the attempt and its error are recorded, which is exactly the case an incident review cares about.
  • The tool set is bounded on purpose. Attributes are keyed by a stable tool name, so a renamed tool shows up as a migration item in the query results rather than silently splitting one tool's history in two.

If your agent framework calls tools over HTTP without a protocol level, the same design applies with different vocabulary: bind the tenant at the edge, re-assert it per call, and record the attempt rather than the success.

AI agent security: tenant isolation is a query property

Observability features fail in a specific and unpleasant way: the dashboard is scoped, the API behind it is not. The defence is unglamorous — make the tenant part of the query builder, so that a query without a team id cannot be expressed. In this codebase the filters are a dataclass that the SQL builder consumes, and the builder is the only place a WHERE clause is assembled:

# backend/smartgate/core/audit_filters.py — source lines 58–69 (AuditQueryFilters)
@dataclass
class AuditQueryFilters:
    tools: Optional[Sequence[str]] = None
    success: Optional[bool] = None
    route: Optional[str] = None
    scope: str = "core"
    query: Optional[str] = None
    since: Optional[str] = None
    correlation_id: Optional[str] = None
    event_kind: Optional[str] = None
    request_id: Optional[str] = None
    trace_id: Optional[str] = None

Note what is not in that dataclass: a team id. The tenant is a separate mandatory argument on every query, so "forgot to filter by team" is a type error rather than a code review finding. The remaining fields are the filters an investigation actually needs — by tool, by outcome, by route, by scope (a stream can ask for its own core subset, or for the full trail), and by a free text query. The same builder serves the interactive query, the trace view and the CSV export, so the three surfaces cannot drift apart in what they are allowed to see.

Two consequences for anyone designing this: an empty or missing team id must raise, not return everything (here it raises AuditQueryError), and the "full" scope should be a deliberate, auditable choice rather than the default, because it is the one that widens what an operator can read.

Agent observability: one trace, many calls

Once every call is a row, the useful unit is still the task — "summarise these six pages" is one intent expressed as six calls. The row-level fields give you the pieces; a correlation key gives you the whole. The trace query is the smallest interesting one: take a trace id, return every row that carries it, oldest first, with the count of what the tenant can see:

# backend/smartgate/core/audit.py — source lines 196–211 (query_by_trace_id)
async def query_by_trace_id(
        self,
        team_id: str,
        trace_id: str,
        *,
        limit: int = 200,
    ) -> Tuple[List[Dict], int]:
        if not team_id or not str(team_id).strip():
            raise AuditQueryError("team_id is required")
        if not trace_id or not str(trace_id).strip():
            raise AuditQueryError("trace_id is required")
        if self._pool is None:
            return [], 0

        flt = AuditQueryFilters(trace_id=trace_id.strip(), scope="full")
        where_sql, args = build_audit_where(team_id, flt)

Three decisions are visible in fourteen lines, and all three are load-bearing. The signature requires team_id and trace_id, so a trace id leaked from a log cannot be used to read another tenant's task. Ordering is ascending, because a trace is read as a story and not as a timeline of latest events. And the count comes from the same WHERE clause as the rows, so a paginated UI cannot claim "412 events" while showing three — the two numbers are computed from one predicate.

The reading that follows is the one engineers usually skip: a trace with two calls where the second never started is a transport failure, while a trace with two successful calls and no third is a planning failure. Both are one query away, and neither is visible in a prompt-level trace at all.

AI agent governance: the export path is the control

Governance work tends to arrive as a request from outside engineering: the security team wants the audit trail in their SIEM, the auditor wants a date range, and both want to know the export is complete and bounded. That is a different read pattern from the UI — ascending by time, capped hard, and shaped so a downstream system can ingest it:

# backend/smartgate/core/audit.py — source lines 227–240 (query_for_export)
async def query_for_export(
        self,
        team_id: str,
        *,
        filters: Optional[AuditQueryFilters] = None,
        limit: int = 10000,
    ) -> List[Dict]:
        """Fetch audit rows for SIEM export (ASC by timestamp, capped)."""
        if not team_id or not str(team_id).strip():
            raise AuditQueryError("team_id is required")
        if self._pool is None:
            return []

        cap = max(1, min(int(limit), 50000))

The cap is the interesting part. It is clamped between 1 and 50,000 rows per call, and the default of 10,000 is a deliberate trade: an export that silently truncates is worse than one that fails, so the caller always receives an explicit number of rows and can walk the window forward. The default filter scope is the narrower core set rather than the full trail, which keeps bulk exports from becoming an accidental disclosure channel, and the whole read runs through the same tenant-scoped builder as everything else. If you are asked for a SIEM export and your design has a separate "admin export" query, that second query is the one that will be wrong in six months.

The retention side of governance is a policy decision, not a code decision: the plans in this product keep audit rows for 7, 30, 90 or 180 days depending on tier, and the retention job runs as a scheduled task with its own authorisation rather than as a side effect of normal traffic. Write the retention period next to the export contract, because "how far back can I prove this" is the question auditors actually ask.

AI agent monitoring: roll up the task, not the call

Per-call success rates are a poor health signal for agent work: a task that is 83% successful at the call level can be 0% successful at the intent level, because the failure lands on the last step. The rollup query aggregates rows into tasks by correlation key and returns the fields a status page needs in one pass:

# backend/smartgate/core/audit.py — source lines 160–173 (query_tasks)
                SELECT
                    params->>'correlation_id' AS correlation_id,
                    MIN(timestamp) AS started_at,
                    MAX(timestamp) AS ended_at,
                    COUNT(*)::int AS span_count,
                    COALESCE(SUM(token_used), 0)::int AS total_tokens,
                    BOOL_AND(success) AS all_success,
                    BOOL_OR(success) AS any_success,
                    array_agg(DISTINCT tool ORDER BY tool) AS tools
                FROM audit_logs
                WHERE {base_where}
                GROUP BY params->>'correlation_id'
                ORDER BY MAX(timestamp) DESC
                LIMIT ${limit_idx} OFFSET ${offset_idx}

This single query answers five questions that would otherwise be five dashboards. span_count is the length of the task. total_tokens is its cost, summed rather than estimated. all_success and any_success together separate "this task worked" from "this task worked eventually" — a retry succeeds, and the pair of booleans keeps that visible instead of flattening it. And array_agg(DISTINCT tool) gives the tool sequence as a set, which is how you notice that a task which used to take three calls is now taking nine. Ordering by MAX(timestamp) descending puts the most recent task on top, which is what an operator opening the page at 09:00 wants.

Langfuse alternative: what prompt-level tracing leaves out

If you already run a prompt-level tracing tool, the honest comparison is not "which dashboard is prettier" but "which events exist in the record". The categories overlap: both can show you a model call's tokens and latency. They diverge on the caller, the route and the transport, which are the facts an access review or an abuse investigation starts from:

# backend/smartgate/core/audit_enrichment.py — source lines 43–47 (client_source_ip)
def client_source_ip(headers: Mapping[str, str]) -> str:
    raw = _header(headers, "x-forwarded-for") or _header(headers, "x-real-ip")
    if not raw:
        return "anonymous"
    return raw.split(",")[0].strip() or "anonymous"

That function is four lines and it is the reason this page exists as a companion to prompt-level tooling rather than a competitor to it — the caller's address is recorded with the call, defaults to anonymous rather than an empty string, and is taken from the forwarding headers an edge actually sets. Pair it with the two derived attributes the same enrichment pass adds (the inferred route and the inferred transport) and each row can answer "who reached this, over what, and through which entry point", which is the question a trace of prompts and completions cannot answer at all.

Question Prompt-level tracing Tool-call record (this design)
What did the model say? Yes, payloads stored Not stored by design
What did the agent do? Only if the model call encloses the tool Yes, one row per call
Which tool spent the money? Estimated per generation Summed per tool and per task
Who called it, from where? Usually absent Source address, route and transport on the row
Can it be exported to a SIEM? Vendor-dependent Bounded, ascending export query
Retention as a policy knob? Vendor-dependent 7 / 30 / 90 / 180 days by plan

The two are complements, and the ordering matters: get the tool-call record right first, because it is the one your cost, security and compliance questions all resolve against.

AI agent cost: per-call token accounting

Cost work on agent systems fails for a mundane reason: the number that arrives is an estimate of a whole request, and the decision you have to make is about a single tool. The fix is to write a token count on the call row at the moment it completes, which is what token_used is — the same value the budget guard reads when it checks a team's monthly allowance, and the same value the task rollup sums. Where an upstream tool does not return an exact count, the gateway falls back to a local estimate rather than leaving the column empty, because an approximate number that exists is comparable week over week, whereas a NULL is not.

Three habits make that column useful. Record tokens and latency on the same row so cost per millisecond of saved work is a query, not a spreadsheet. Keep the token count for failed calls too: work that was paid for and then thrown away is precisely the spend a budget review is looking for. And keep the counter monotonic per call — retries are separate rows, never an in-place addition, so "we paid twice" stays legible afterwards.

How to get started

If you have no observability at all, start smaller than a platform. Write the eight fields for one tool call to wherever you already log, and check that you can answer two questions a week later: how many calls failed yesterday, and which tool used the most tokens. If either answer requires reading source, the record is missing a field rather than a dashboard. The audit and compliance pillar describes the product surface this design feeds; the MCP-specific logging questions live on mcp logging and observability, and the gateway-level picture is on mcp gateway.

To see the record with real calls flowing through it, start free — the free plan keeps seven days of audit rows for every tool call, which is enough to run the two questions above against production traffic. Compare the retention and export terms against your own requirements on the pricing page before you commit a compliance deadline to them.

Frequently Asked Questions

Limitations and what this does not do

The record described here is a tool-call record, and it is honest about its blind spots. It does not store prompt or completion text, so it cannot answer "what exactly did the model see"; that is a payload question with different retention and privacy constraints, handled separately. It cannot detect a bad answer — a call that succeeded with an unhelpful result is success: true, and quality evaluation is a different discipline with different data. The default shape favours completeness over volume: one row per call, no sampling, which is cheap per row but adds up on a high-volume plan, and teams that log every call should budget storage as deliberately as they budget tokens. Where a token count comes from a local estimate rather than an upstream tool, it is an estimate and should be treated as one. And none of this replaces an access review: an audit trail shows what happened, but someone still has to decide whether it was allowed.

Sources

Method note

The code excerpts in this article are not transcribed. Each block was cut directly out of the slice body returned by the SmartGate slice API and re-asserted byte-for-byte as a substring of that body before publication; the first line inside every fence records the file and the exact source lines it came from. Symbols were pinned by a whole-name match confirmed by a server-side proof call before any of them entered the text, and the two sections that could not be pinned uniquely are written from the product's documented behaviour and from the external sources listed above — a section that cannot be pinned is sourced, never invented. Repository-relative paths are shown as they are in the source; internal hosts, credentials and private addresses were stripped before any of it reached this document.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)
1 llm observability AuditLogger backend/smartgate/core/audit.py 39–49 rule A L2 → slot-proof 8abeddeef9a4
2 ai agent security AuditQueryFilters backend/smartgate/core/audit_filters.py 58–69 rule A L2 → slot-proof 9f79bd88a9b0
3 agent observability query_by_trace_id backend/smartgate/core/audit.py 196–211 rule A L2 → slot-proof be3d20fdc94a
4 ai agent governance query_for_export backend/smartgate/core/audit.py 227–240 rule A L2 → slot-proof 2c40e537d29f
5 ai agent monitoring query_tasks backend/smartgate/core/audit.py 160–173 rule A L2 → slot-proof 35caff93e2b9
6 langfuse alternative client_source_ip backend/smartgate/core/audit_enrichment.py 43–47 rule A L2 → slot-proof 8bf9501b9e29

Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before publication. 6 of 8 sections pinned, 0 abstentions, 2 misses.