SmartGate

AI Compliance: The Evidence Trail an Auditor Expects

AI compliance is not a document you publish but a record you can produce on demand: one row per model or tool call carrying who asked, which key was used, which tool ran, when, and how it ended; a retention period taken from the longest duty that applies to you rather than from a default; an export a third party can read and check without your help; and documentation that states what the system…

Short answer: AI compliance is not a document you publish but a record you can produce on demand: one row per model or tool call carrying who asked, which key was used, which tool ran, when, and how it ended; a retention period taken from the longest duty that applies to you rather than from a default; an export a third party can read and check without your help; and documentation that states what the system records instead of asserting that it makes you compliant.

ai compliance: the four questions an evidence review asks

A requirement gets easier once it is translated into a mechanical question: can you produce the record? A reviewer does not want proof that a policy exists in a document store; they want the row for one particular call — who triggered it, with which credential, against which system, when, with what result. Four tests follow, and none of them is a legal judgement:

  • Completeness. Does the record cover every call in the class you claim to log, or only the ones that survived a sampling pipeline tuned for cost?
  • Attribution. Can a reader name a responsible party per row — a person, a team, a workload — rather than a shared service account that everything authenticates as?
  • Retention. Is the period derived from an obligation that applies to you, or copied from a product default?
  • Integrity and export. Can the record leave your infrastructure unmodified, readable and checkable by someone else?

Observability alone is not evidence. A tracing product is built to answer "what is happening now" cheaply, so it samples, aggregates, truncates payloads and deletes on a short schedule — all correct for operations, all wrong for a record. A sample proves nothing about the calls it did not keep, an aggregate cannot name a caller, and a truncated payload cannot show what a tool was asked to do. Two stores over one event is the shape that works.

The word in the heading is also the first trap: "compliant" is a conclusion about a specific system used in a specific way by a specific organisation, and no vendor reaches it on your behalf. What can be stated and tested is narrower — these fields are recorded, this is the retention, this is the export, and this is the clause each artefact is aimed at.

ai governance compliance: the frameworks, and the artefact each one names

Frameworks read as competing checklists, but they name a small number of artefacts and mostly agree on them. The table maps the clause a reviewer is likely to quote to the thing that answers it; read it as a coverage test rather than a conformance verdict.

Clause or control The artefact that answers it
EU AI Act Art. 12(1)–(2) An event log the system writes itself, over the system's lifetime, covering the events behind risk situations, substantial modification and post-market monitoring
EU AI Act Art. 19, Art. 26(6) A retention period in days plus the deletion job, matching the six-month floor that binds provider and deployer alike
EU AI Act Art. 11, Annex IV, Art. 13 Per-system documentation of what is logged, in what format, for how long
NIST AI RMF GOVERN 1.1, MANAGE 4.1 An obligation register, plus a post-deployment monitoring plan naming events, thresholds and alert owners
ISO/IEC 42001 clause 9.1, Annex A.6.2.6, A.6.2.8 Evidence that monitoring ran, with results and dates, plus the retained event log itself
SOC 2 CC6.1, CC7.1–CC7.4 Access records per credential, a change history for the tool and model inventory, an alerts log with resolution notes
ISO/IEC 27002 controls 8.15, 8.16 Log sources defined, protected and analysed on a stated schedule
GDPR Art. 30, Art. 5(1)(e) A retention rule per field, and a lawful basis for keeping prompt content at all

The dates attached to the Act matter more than most summaries suggest. The Article 5 prohibitions and the Article 4 literacy duty have applied since 2 February 2025; the Chapter V general-purpose model obligations since 2 August 2025; the Article 50 transparency duties since 2 August 2026. Regulation (EU) 2026/1744, the Digital Omnibus on AI, then deferred the Annex III high-risk regime to 2 December 2027 and the Annex I regime to 2 August 2028 — which moves the deadline, not the design question. A record cannot be created retroactively, so an obligation starting in December 2027 asks for a log that was already running: preparation starts when the system first handles real traffic.

One well-shaped audit row serves Art. 12, Annex A.6.2.8, CC7.2 and 8.15 at once, so build one record and map it rather than four. And the retention numbers disagree — a six-month floor, a fiscal year, a three-year contract, a ninety-day plan tier — so the record has to outlive the longest of them.

ai audit trail: the fields one row must carry

An audit trail is a schema before it is a system. These fields are the ones that survive contact with a reviewer, ordered by how often their absence is the finding.

Field What breaks without it
Event id, UTC start and end Two calls in one second become indistinguishable; a timeout reads as a refusal
Principal A shared service account answers what happened and never who did it
Credential id, team, plan You cannot revoke the key behind an incident, or scope the row
Transport, endpoint, method You cannot show which interface was used or which policy applied
Tool and server identity A tool ran is not an answer when four servers expose one name
Outcome, error class, policy rule id If only successes are written, refusals vanish; a denial with no rule id is an unexplained gap
Model, token counts Cost questions get answered from provider invoices instead of the record
Latency, retry index Duplicate rows look like duplicate actions
Correlation id One user action across three services cannot be reassembled
Request digest, not request You either keep content you cannot justify, or cannot show what was sent

Four disciplines turn that list into something that survives an audit. Use UTC in one unambiguous format, and record start and end rather than a single timestamp, because duration is how a reviewer separates a timeout from a refusal. Write one row per attempt, not per success: a control that denies a call is the most valuable evidence the system produces. Never write the credential itself, only its identifier. Version the schema, because an export is read years later, when nobody remembers that the outcome field once held free text.

Publishing a versioned contract that outlives the code generating it is the practice set on API documentation best practices for AI tools.

The field most often omitted on purpose is the payload. Keeping a digest rather than the content proves that a specific input produced a specific call without retaining the input. Where content is genuinely needed, treat it as a separate record class with its own retention, access list and lawful basis: prompt text is personal data the moment a person is in it.

ai agent audit: attributing a call to an agent, not an account

A single request can pass through a session, an orchestration run, several steps and a tool call before anything leaves your network, and every hop has a natural identity. A record that keeps only the last hop answers what happened and fails the reviewer's real question, which is who caused it.

Two fields carry the weight. An explicit on-behalf-of attribute carries the human or system that initiated the work alongside the credential that executed it; where an agent vends its own key, the key identifies the agent and the attribute identifies the principal, and only the pair supports an access review. An agent identity that is first-class in the row — an id plus a version, since the instruction set, the tool set and the model behind that identity are all part of what acted — keeps last quarter's decision explainable this quarter.

Three consequences follow. Delegation belongs in the record: if a supervisor hands work to a worker, show both, so a reviewer can tell a wide action by one agent from a narrow action by three. Retries must be distinguishable from repeats, so the idempotency key goes in the row and not only in the application. And where one call is visible from several places — the agent host, the tool boundary, the model provider — choose an authoritative source and correlate the rest by trace id.

There is a jurisdictional reason for per-agent clarity too: where a tool is used inside a regulated activity, retention duties usually attach to the regulated institution, which must show which authorised actor used which capability. That is the same record the AI agent workflows for financial data analysis discussion needs in order to be defensible: the duty is discharged by showing who authorised what, when, not by a model choice.

mcp logging: what the protocol gave you, and what changed

An MCP deployment once had a protocol-native answer here. The specification's Logging utility let a server push structured log messages to its client as notifications, with severity levels negotiated through a set-level request and the capability declared at initialisation. Be blunt about what that was: a debugging channel aimed at the client, whose copy the client decides whether to keep — never a retention mechanism, and never under the server operator's control.

Its status has changed. In the 2026-07-28 revision of the specification, the Logging feature is deprecated and new implementations are advised not to adopt it; the guidance is to write structured logs to the server's standard error stream, or to emit telemetry through OpenTelemetry. The same revision documents how trace context propagates through the protocol's metadata fields, so a request can be followed across a client and the servers it reaches.

The compliance consequence is the part worth internalising: once the log is no longer a protocol message, completeness, protection and retention are properties of your deployment and not of the specification. Nothing in the protocol promises that any given log line is written, stored or kept for six months. That makes the tool boundary the natural recording point — the place that sees the call, the key and the outcome in one object. How that boundary is structured is the subject of the MCP gateway page; from the evidence side, what matters is that the row is emitted where the action happens, not reconstructed later.

What that boundary looks like from a person's side — one server block in an assistant they already run, with every call landing in the same trail — is on SmartGate MCP for research and decisions.

One administrative trap to close: the protocol's telemetry attribute names have moved between semantic-convention registries and now sit with the generative-AI conventions. Pin the version you emit, because an export whose field names drift between revisions stops being comparable.

ai agent monitoring: from an event stream to a decision

Monitoring is what you do with the rows, and it is where teams spend the most effort for the least evidence. Alert on a short list of states that each map to a named control:

  • Refusal rate per key and per tool. A denial is a control working; a spike is an over-tight policy, a misconfigured client, or an attempt a human should read.
  • Request rate against that key's own ceiling. Per-key limits differ by plan, so the alert should name the key approaching its limit rather than a global number.
  • Error classes per tool. A new error class usually means the tool changed underneath you, which is a change-management event and not only an engineering one.
  • Token burn against the monthly cap. It turns a budget conversation into a control event with a row attached.
  • A new tool, a new server, or a changed tool description. The inventory diff is the highest-value alert in an AI deployment, because a widened capability is what most reviews are trying to reconstruct afterwards.

Two rules stop the alerting layer being mistaken for the record. Monitoring and evidence are separate paths over one event: the alerting path may sample, aggregate and drop, the retained path may not. And an alert is not the evidence — the resolution note is, so the alert log needs the same retention discipline as the audit rows. The operational side of this, and what a trace can and cannot prove, belongs to LLM observability; this page's interest stops at the boundary between a stream you watch and a record you keep.

agent observability: the minimum signal set that survives scrutiny

Agent observability has converged on one shape: a span per model call and a span per tool call, nested under the request that caused them, with attributes rather than prose in the span name. That shape suits the evidence job, because a span is addressable — it has a parent, a duration and an identity that survives being copied elsewhere.

Two properties of the usual implementation make it insufficient on its own. Sampling is the right economic decision for a high-volume service and the wrong basis for a claim about what was recorded: if you say you log every tool call, a few-percent sample cannot be the only copy. Sampling is legitimate for the traffic you do not claim — development keys, internal experiments — and evidence needs complete coverage of the class you do claim. Time-to-live is the second: trace stores typically keep data for hours to days, longer than a debugging session and far shorter than a retention obligation, so the two stores need different lifecycles, budgets and access lists.

The minimum set worth calling an evidence signal is small: an identifier per event, the actor and credential, the capability that ran, the outcome, the timing, and a correlation id that ties the chain together. That is the list the audit-trail section specified, which is the point — observability and audit are the same event kept for two lifetimes. Define the event once with a stable, versioned schema, route it to a short-lived operational store and a long-lived retained one, and treat the retained copy as the source of truth.

mcp monitoring: health, limits and drift at the tool boundary

Where a tool boundary exists, it sees what no client sees, and several of those things are control events rather than performance metrics.

Inventory. Which servers are registered, which tools each one exposes, what each tool's description and input schema say today, and when that last changed. A tool description is part of the instruction the model reads, so a rewritten description changes system behaviour and a changed schema changes what the system can be asked to do. Detecting that diff is the operational form of detection of configuration changes, and it is cheaper than reconstructing a change after an incident.

What a rewritten description or a swapped definition can actually do to a running agent — tool poisoning in the rug-pull variant — is the runtime surface on AI agent security.

Limit events. A request rejected against the key's per-minute allowance, a denial because a tool is not permitted for that key, a plan ceiling reached mid-month: each is a control firing, and each belongs in the retained record with the key and the tool named. Where the boundary enforces the per-key and per-team ceilings, the audit row and the control decision come from one object, which is what makes the record explainable to someone who was not there.

Which layer owns those ceilings, and how over-reach is noticed while it is happening rather than afterwards, is the control plane on AI agent governance.

Key lifecycle. Keys are created, rotated and revoked, and the record has to survive all three: attribution must resolve through a stable key identifier, because once a key is deleted its readable name resolves to nothing. Revocation should appear as an event in the trail, not as an absence of traffic afterwards.

The ceilings are operational facts, not compliance advice, and the plan table is authoritative:

Plan fact Tier values Why a record cares
Monthly token ceiling 2M / 20M / 100M / 200M+ Bounds the volume the row stream can reach
Requests per minute per key 120 / 300 / 600 / 1200 Sets the rate limit events are written and alerted on
Audit-log retention 7 / 30 / 90 / 180 days Ceiling on readability before the record must be exported
Keys per team 2 / 10 / 30 / 9999 How finely a team can attribute calls to distinct credentials

How to export evidence for an auditor

An export is the point at which the record becomes someone else's data.

  • A stable, versioned schema. One row per event, one field name per meaning, UTC timestamps, in a format the reviewer's tools can open — a format that needs your own product is not an export.
  • A manifest. Row count, window start and end, the filter applied, the schema version and a hash of the file. Without one, the recipient cannot tell whether the rows they hold are all the rows.
  • A redaction decision made before the export. Which fields are personal data, which are secret-adjacent, what the recipient may hold. An export to an external audit team is a disclosure with a purpose, and the safe default is the metadata row with digests in place of content.
  • A schedule shorter than the retention ceiling. A record kept 180 days on the top tier is unavailable on day 181 unless someone exported it, so export on a cadence shorter than the shortest duty you carry and keep the copy under your own policy. Where the duty is a six-month floor, note that 180 days of calendar retention is the closest tier to it and still a plan limit, not a legal opinion — read the current pricing page and check the arithmetic against your own deadline.
  • Chain of custody. Who exported what, when, to whom, and which rows were filtered out. The recipient will ask why the manifest count is lower than the traffic they expected, and a documented filter is the answer.

Design against one test: an auditor should be able to reconstruct a single decision end to end — request, credential, tool, result, effect — without a call with your engineers. If that needs a walkthrough, the export is documentation about your system rather than evidence from it.

What you may claim, and what you may not

Most compliance problems in tooling documentation are claim problems rather than engineering problems.

Safe, because testable: this event carries these fields; this field is retained this many days on this plan and deleted by this job; the record exports in this format with this manifest; this artefact is what a review asks for under this clause, for a system of this kind.

To avoid: that a product or a plan "satisfies" a regulation, that a tool "makes you compliant", that a deployment is "audit-proof", that anything is "guaranteed" — and the twin of that last one, presenting an attestation report as a certification. A service-organisation report covers a service organisation's controls over its own system, as assessed by a service auditor: it is not a certificate for a product, and it says nothing about the controls of the customer running it.

Two boundaries do most of the work. The actor boundary: obligations attach to the role an organisation plays — broadly, the party that puts a system on the market and the party that uses it under its own authority — so one tool demands different evidence from a provider and from a deployer, and different evidence again from a deployer who starts acting as a provider by putting its own name on the system. The data boundary: retained prompt or completion content is personal data, so "we keep everything for six months" is a privacy decision with a lawful basis behind it, not a compliance feature in front of it, and a retention period chosen for a logging duty can conflict with a storage-limitation duty. Where content must be kept, give it its own rule and keep the metadata record separable from it. What may enter a prompt at all is the input-side question owned by the secure prompt handling in AI applications page.

Frequently Asked Questions

Limitations

This page describes what a record must contain, how a retention period is chosen and how an export is produced. It is not legal advice, and it does not determine whether any system is high-risk, which obligations attach to your organisation, or how long you must keep a given record — those depend on the system, the use, the jurisdiction and the contracts; the current text of each instrument is the only authority for them.

Nothing here claims that any product, plan or configuration satisfies a regulation, a framework or an audit. The plan figures quoted are operational limits read from the published plan table, not a statement about a duty: a retention tier is a technical ceiling on readability, and the compliance question is whether that ceiling is longer than the duty that applies to you.

The mapping table is a coverage aid rather than a conformance verdict: control identifiers and clause numbers are quoted so a reader can find the source text, and a control is satisfied by an operating practice over time, not by the existence of a field. Framework texts are revised as well — the Act has already been amended once and its high-risk dates moved — so verify the current text before relying on a date on this page.

This page carries no code excerpt, and the reason is recorded in the Method note below: the slice matcher found no unique symbol for any of its eight sections. No line-numbered claims, no quoted implementation detail and no assertion about the internal shape of any product named here follow from that.

Sources

  • Regulation (EU) 2024/1689 (AI Act) — Article 12, record-keeping, Article 19, automatically generated logs, Article 26, obligations of deployers and Article 72, post-market monitoring, read on 2026-09-30 for the six-month retention floor.
  • Regulation (EU) 2026/1744 (the Digital Omnibus on AI), published in the Official Journal in July 2026, for the amended application dates: the Annex III high-risk regime moved to 2 December 2027 and the Annex I regime to 2 August 2028. The European Commission's implementation timeline reflects the amendment.
  • NIST AI Risk Management Framework 1.0 (NIST AI 100-1), for GOVERN 1.1 and MANAGE 4.1 — the framework core.
  • ISO/IEC 42001:2023, for clause 9.1 and Annex A controls A.6.2.6 and A.6.2.8 — iso.org.
  • AICPA Trust Services Criteria (2017, revised 2022), for CC6.1 and the CC7.1–CC7.4 monitoring and incident-response series this page's alert set is mapped to; ISO/IEC 27002:2022, controls 8.15 (Logging) and 8.16 (Monitoring activities).
  • The Model Context Protocol specification, revision 2026-07-28, for the deprecated Logging feature and the guidance to write to the server's standard error stream or emit telemetry through OpenTelemetry, plus the trace-context propagation in _meta; the OpenTelemetry semantic conventions, where the protocol's attributes now sit with the generative-AI conventions.
  • Demand figures in this page are our own measurements: DataForSEO Google Ads, United States, 12-month window, measured 2026-09-30, recorded in this project's search_volume.json and research_brief.md.

Method note

This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher recorded 0 of 8 sections pinned for this page (0 abstention(s), 8 no-slice verdict(s)): for each of the eight section keywords, rule A found no unique symbol in the scanned repository — this lane's vocabulary (compliance, audit, logging, monitoring) collides with generic names across a product codebase. A pinned generic would have given the page the shape of a verified article with none of the substance, so every section above is written from the sources listed.

Product claims were read from the product source at the revision the slice run recorded in this project's pipeline_results.json, read-only, and the plan figures were re-verified against the live pricing page on 2026-09-30. Clause numbers, control identifiers and application dates were re-read on that date from the instruments listed under Sources. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.