SmartGate

AI Compliance Certification: Materials, Windows, Findings

AI compliance certification is an audit of a management system, not of a model: an accredited body reads your process against a scheme's clauses, samples the records that process produced, and issues a certificate whose scope names what it covers.

Short answer: AI compliance certification is an audit of a management system, not of a model: an accredited body reads your process against a scheme's clauses, samples the records that process produced, and issues a certificate whose scope names what it covers. The work runs through four phases — scope and gap assessment, an operating period that generates evidence, internal audit and management review, then a two-stage external audit — and the date is set by the earlier of two clocks: when an obligation starts applying, and when a customer asks to see the certificate.

Key takeaways

  • The auditable unit is the process. Nothing in a certificate says a model is accurate; it says a management system ran, and a sample proves it.
  • Evidence cannot be back-filled. The auditor samples a past period, so records must start before the visit is booked — the operating period, not the audit, is the long pole.
  • Internal audit and management review are read first. A missing decision in the minutes is a finding no other artefact repairs.
  • Most findings are process findings: scope statements describing a different system, risk assessments never revisited, actions closed without a root cause.

ai compliance certification: four phases between the gap assessment and the certificate

A certification engagement has the shape management-system certification has had for two decades: a scheme with auditable clauses, a body accredited to audit against it, and an auditor who samples what the process produced. The object under audit is your process. No one asks a model to justify an output; someone asks what the procedure says, who followed it, and what the record of that looks like.

Scope and gap assessment. Fix what is inside the management system — which AI systems, sites, teams and suppliers — then read the scheme's clauses against what the organisation does today and write every gap with an owner and a date. The phase ends with a scope statement, a gap report and a first evidence register mapping clause to document. Scoping wider than you can operate is the expensive mistake here: every added system owes a risk assessment, an owner and records that stay current.

An operating period. The scheme wants evidence that the process ran, so it has to run long enough to leave a sample: a quarterly review needs a quarter, an annual internal audit a year. The phase carries no external deadline, which is why it slips, and it decides the whole schedule.

Internal audit and management review. These two the external auditor necessarily reads, because they are the system's own checks on itself. The internal audit has to be independent of the area it audits, and the review has to produce decisions with owners and dates rather than a status update.

Stage one and stage two. Stage one is a documentation and readiness review — is the system described well enough to audit. Stage two is the implementation audit, where the sample is taken. Findings follow, and a major one can require a follow-up visit before the certificate issues.

Phase What exists at the end Lead time to allow Who owns it
Scope and gap assessment Scope statement, gap report, evidence register Weeks Whoever will own the management system
Operating period A period of records nobody assembled for a visitor Months — one cycle of the shortest recurring activity Process owners, one per clause area
Internal audit and review Findings plus minutes carrying decisions Weeks, on a calendar An auditor independent of the area audited
Stage 1 and stage 2 Audit report, findings, corrective-action plan, certificate Weeks to months, plus any follow-up The certification body, with a client-side coordinator

Where the request-time decisions inside the scope sit — what a policy does with an incoming payload, how a key is validated — is the mechanics on secure prompt handling in AI applications. This page starts one level up: which decisions are written down, and what can be sampled.

ai governance certification: choosing a scheme and a certification body

The phrase covers three different transactions, and only one ends with a certificate for a procurement pack. A person-level credential is a training outcome. A vendor's service-organisation report is an attestation about that vendor's own controls. A management-system certificate is issued to the organisation that runs the system, by an accredited body against a scheme — and it alone carries a scope statement a customer can read.

The scheme usually has two candidates. The management-system route is ISO/IEC 42001:2023, whose clauses four to ten carry the system itself (context, leadership, planning, support, operation, performance evaluation, improvement) and whose Annex A lists the AI-specific controls. An existing information-security certificate is worth mapping rather than chasing: the real question is which one the customer asked for.

The body is where teams under-invest. Four checks do most of the work.

  • Accreditation and its scope. The body should be accredited for that scheme, and the scope readable — a certificate from a body whose accreditation does not reach the scheme is a document without a backer.
  • Competence of the team. ISO 19011 makes auditor competence a requirement, and for an AI system the team has to be able to read a risk assessment about a model, not only about a server. With several locations in scope, the sampling plan decides whether one site carries most of the effort. The pack a body asks for at quotation time is short and predictable: scope statement, policy set, risk register, the management system's organisation chart, and a note on which AI systems are live. Producing it honestly is a better readiness test than any self-assessment, because the body prices the audit from what it finds. What a certificate proves once you hold it is the separate question on AI governance certification; here the concern is only to get the scheme and the auditor right, since both are hard to change mid-cycle.

ai compliance framework: which documents can be certified against

The word "framework" does the damage in this query: the instruments people collect under it are not the same kind of object, and only one kind carries a certificate. Sort them by what happens when you ignore one.

Instrument What it is Who checks it Certificate available
Regulation (EU) 2024/1689 (AI Act) Law Market-surveillance authorities No — a legal duty is not audited into a certificate
NIST AI RMF 1.0 Voluntary guidance Nobody on your behalf; you measure yourself No
ISO/IEC 42001:2023 Certifiable management-system scheme An accredited certification body Yes, with a scope statement
A service-organisation report (SOC 2 / ISAE 3402) Attestation about one organisation's own controls A service auditor No — the report is the deliverable
A customer's AI questionnaire Contract The customer Not applicable

Read the table in one direction and the preparation plan falls out. The law says what must be true; the guidance says how to decide and how to document the decision; the scheme is what a third party audits and the only row ending in a document with your name on it. The questionnaire is usually what starts the exercise — worth knowing before choosing a scheme, because if three customers ask three frameworks, one certificate should answer the most of them.

Two habits keep the document set maintainable, and both are about not mixing classes. Keep one register mapping requirement to artefact to owner to last-verified date, so a change of regime moves rows instead of rewriting prose. And hold the scheme's distinction between documented information you maintain and records the process produced: the first can be edited to match reality, the second cannot be edited without ceasing to be evidence.

The mapping from obligation to role, and role to artefact, is set out on AI compliance framework. This page takes it as given and asks the next question: which of those artefacts a certification body actually reads, and in what order.

ai governance compliance: scheduling the audit window

Two clocks decide the date and they rarely agree. The legal clock is when an obligation starts applying to a system like yours; the certification clock is the engagement itself — an operating period, an internal audit cycle, two stages of external audit, and the corrective-action window after findings. Work back from the earlier deadline, with the operating period in front of the audit.

The legal clock has four dates worth holding. The prohibitions and the AI-literacy duty have applied since 2 February 2025; the general-purpose model obligations since 2 August 2025; the Article 50 transparency duties since 2 August 2026, already in force at the time of writing. The high-risk regime moved to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I under the Digital Omnibus amendment. A record cannot be created retroactively, so the date that matters is not when an obligation applies but the date you had to be operating to have a sample for it.

The certification clock is duller. Stage one precedes stage two; surveillance audits follow at intervals the scheme defines, commonly within twelve months; recertification is a fresh audit before expiry, on the usual three-year cycle. The period the body samples runs from the previous visit to the closing meeting, which is why a certificate improvised in the last month produces a first surveillance visit that reads like a full audit.

You need the certificate by Start the gap assessment Have the process running by Book stage 2 for
A bid deadline in nine months Now Within one month Months 7–8, leaving room for a follow-up
An obligation starting in eighteen months Within a quarter Within two quarters Roughly a year out
Three scheduling rules survive a real engagement. Put the internal audit in the same month every year,
so the second cycle is not the first repeated under pressure. Leave a quarter between the internal audit
and the certification audit, because your own findings have to be closed before someone else reads them.
And do not schedule stage two until the records span a representative period — an auditor sampling a
quiet month will ask for a busy one instead.

ai audit trail: the records that must exist before the first audit day

An audit is a sampling exercise over a past period, so the honest question is not "can we produce records" but "how far back can we". That window is an entitlement before it is a policy: on our plans the retained audit-log window is 7, 30, 90 or 180 days by tier, with /pricing authoritative. Whatever the ceiling, an auditor asking in September for March is asking whether anything was exported in between.

The disciplines that make the trail audit-ready are procedural.

  • Export on a cadence shorter than the audit period. Keeping the record in the product gives you a window; exporting on a schedule gives you a history, and the export is what you hand over.
  • Test the export before someone else does. Open last year's file, check its manifest against the row count, and record who ran the test. An export nobody has opened is an assumption.
  • Name an owner per record class. A retention decision with no owner is revisited only when a deadline appears.
  • Write the scope statement. Which calls are recorded, over which paths, for how long, exported how often. Without it, a deliberate exclusion and a gap look identical to a reader.
  • Do not repair history. A record rewritten to fill a hole is a new record, and once provenance is lost the sample it was meant to support is worth less than the gap was.

The field-level design — which columns a row carries, how retention is chosen against a statutory floor, how an export is handed over — is the evidence discipline on AI compliance and is not repeated here. From the certification side the requirement is narrower and harder: the trail must be complete across the period the auditor picks, not merely good on the day they arrive.

mcp logging: what an auditor samples at the tool boundary

For an agent deployment the tool boundary is where a sample can actually be pulled, because it is the one place that sees the call, the key and the outcome in a single object. That makes it the recording point an auditor asks about first, and the place where "show me twenty calls from March" is either a query or a project.

The protocol no longer offers an answer here. The Model Context Protocol's Logging utility — a server pushing structured log messages to its client — is deprecated in the specification revision 2026-07-28, with new implementations advised to write structured logs to standard error or emit telemetry through OpenTelemetry. Once the log is no longer a protocol message, completeness, protection and retention are properties of your deployment, and nothing in the protocol is evidence that a line was written or kept.

Three things follow. Capture the tool and server identity with the call, because a tool name is ambiguous the moment two servers expose the same one and a reconstructed decision has to name which server answered. Keep the inventory as it changed — which tools each server exposed, when a description changed — because calls against an unstated tool list describe an unstated system. And keep the sample reproducible: an auditor who asks for one call and gets a working query is finished in ten minutes, whereas one given a dashboard walkthrough records what they were shown rather than what they checked.

Which credential may call which tool, and where a limit is enforced rather than described, is the control plane on AI agent governance. The certification question is the record of it: whether the sampled calls attribute to a credential that existed, under a policy in force, on the date the row carries.

agent observability: the readings your management review must show

The scheme requires the management system to monitor and measure what it does, and the review to use those results. That pairing is where a good tracing stack can still fail: the clause does not ask for a dashboard, it asks for a reading with a date, a threshold and somebody who acted on it.

The evidence is small and specific. A monitoring plan naming what is measured, how often and against which threshold — what NIST AI RMF 1.0 MANAGE 4.1 calls a post-deployment monitoring plan with events, thresholds and alert owners. A periodic record of the readings, retained longer than the audit period. And minutes in which the readings appear as inputs and decisions as outputs, each with an owner and a date.

Two properties of ordinary tooling work against that. Sampling is the right economic decision for a high-volume service and not a basis for a claim about what was recorded, so a monitoring plan saying every call is measured has to agree with the retention behind it. And trace stores usually keep data for hours to days, shorter than any review cycle, so the figures that reach the review must be aggregated on a schedule that outlives the store. What a trace can prove is a subject of its own; the auditor's question is narrower — did the figures in the minutes come from a measurement the organisation actually runs.

ai agent monitoring: continuous evidence across the surveillance year

Once the certificate is issued the job changes shape: the next visit samples everything since the last one, so monitoring keeps evidence continuous rather than producing a good fortnight in October. The failure that prevents is audit theatre — the records dense for two weeks before a visit and thin either side of it.

Cadence Activity The record it leaves
Continuous Calls written to the retained trail; alerts routed to an owner The trail, plus alert resolutions
Monthly Readings against the monitoring plan: refusals, limits, spend, error classes, new tool names The reading set the review later cites
Monthly Check that the retention and export jobs ran A dated job result under an owner's name
Quarterly Revisit open risks, system changes, the tool inventory Updated register entries with a review date
Annually Internal audit, management review, competence refresher Findings, minutes, training records
On material change Re-run the assessment before the change goes live A dated assessment tied to the change

Change is the trigger that catches organisations out, because the cycle assumes the system stays comparable between reviews. A new model behind an existing agent, a new MCP server on an existing key, a new jurisdiction the product is sold into, a rewritten tool description that changes what the agent chooses — each alters the system the certificate describes, and each is cheaper to record when it happens than to reconstruct before a visit. The internal audit programme, with the independence the scheme requires, is the mechanism for checking that this happens rather than asserting it.

Common nonconformities in an AI management system audit

Each finding is generated by a requirement, so reading the requirements as a list of things to demonstrate predicts most of the report. The list is preventive — every row below can be closed by a process change months before an auditor arrives.

Finding The requirement behind it The process fix
The scope lists systems no longer run, or omits one that is The scope has to state what the system covers Re-read it against the production inventory each quarter
Risk assessment unchanged since the last review, despite a new model or tool Risk assessment is a repeatable process, not a one-off document Trigger a reassessment on material change and date the entry
Internal audit performed by the owner of the area audited The audit programme requires independence Split the programme across roles, or use an external auditor
Review minutes report status but take no decision The review has to produce actions Require every input to leave with a decision, owner and date
Competence records missing for the people who operate or review the system Competence must be evidenced, not assumed Keep a record per role, refreshed when the tooling changes
Model providers and MCP servers outside the supplier controls Third parties in the chain are in scope Add them to the register with an owner and a review date
Corrective action closed with a note that the team will be more careful Corrective action has to address the cause Require a root-cause line and a verification date before closure
Records deleted before the audit period closed Retention has to outlive the duty it answers Set export cadence against the audit cycle, not the plan ceiling

Two conventions decide how bad each one is. Grading is usually major or minor: a major means a requirement is not implemented or a systemic failure exists, a minor an isolated lapse, and observations ask for improvement without blocking anything. The corrective-action clock matters as much as the finding — a major is typically addressed before the certificate issues, with a follow-up to verify it, while a minor is verified at the next surveillance visit.

What a gateway changes in the audit preparation

An audit wants records covering a period, so the operational question is where they come from and how long they stay readable. When agent traffic passes through a gateway that holds the credential, the enforcement decision and the record of it come from one component: the call resolves to a key, a team and a plan, is checked against that key's limits, and is written down as it happens. That closes the recurring finding here — a logging claim describing a component nobody deployed.

The plan table is what to check against your own schedule, and these are ceilings rather than compliance statements. Monthly token caps are 2M, 20M, 100M and 200M+. MCP requests per minute per key are 120, 300, 600 and 1200. Audit-log retention is 7, 30, 90 or 180 days — the number that decides how often an export must run to survive an audit period. Keys per team are 2, 10, 30 or 9999, which is the granularity at which a call attributes to a credential rather than a team. Read the current /pricing page before writing any of those into a plan, since it is the authoritative table.

The same record, seen by the person using an agent rather than the auditor sampling it, is described on SmartGate MCP for research and decisions. Two limits belong in the preparation: a gateway records the traffic that passes through it, so a path that never reaches it is absent from every sample, and it does not decide what your process is — the scope statement, the risk register and the review cadence are yours to set.

Frequently Asked Questions

Limitations

This page describes the shape of a certification engagement: the phases, the material, the schedule and the finding classes a clause-by-clause reading predicts. It is not legal advice, it does not determine whether your systems fall under a regime, and it recommends no scheme or body.

The finding classes are derived from the requirements the schemes state, not from a measured sample of audit reports; no statistic about how often a finding occurs is claimed, and grading is governed by the body's own rules. Framework and legal texts are revised — the Act has been amended once and its high-risk dates moved — so verify the current text before relying on a date here. Nothing here claims that any product, plan or configuration satisfies a regulation or holds a certificate; the plan figures are operational limits read from the published plan table, with /pricing authoritative.

This page carries no code excerpt, and the reason is in the Method note below: the slice matcher pinned none of its eight sections. No line-numbered claim, no quoted implementation detail and no assertion about the internal shape of any product named here follow from that.

Sources

  • ISO/IEC 42001:2023, AI management systems — iso.org/standard/42001, for clauses 4–10 and the Annex A controls a certification body audits against.
  • ISO/IEC 17021-1, requirements for bodies auditing and certifying management systems, and ISO 19011, guidelines for auditing management systems — the source of the two-stage shape, the surveillance and recertification cycle, nonconformity handling and auditor competence.
  • NIST — AI Risk Management Framework 1.0 (NIST AI 100-1), for GOVERN 1.1 and the MANAGE 4.1 post-deployment monitoring plan no body certifies.
  • Regulation (EU) 2024/1689 (AI Act) and Regulation (EU) 2026/1744 (the Digital Omnibus on AI), for the application dates scheduled against above — the Commission's implementation timeline reflects the amendment.
  • Model Context Protocol — specification, revision 2026-07-28, for the deprecated Logging feature and the guidance to write to the server's standard error stream or emit telemetry through OpenTelemetry.
  • AICPA Trust Services Criteria, for the attestation route that is not a certification; SmartGate — plan limits read from smartgate.network/pricing.
  • Demand figures are our own measurements: DataForSEO Google Ads, United States, 12-month window, measured 2026-09-30, recorded in this project's search_volume.json.

Method note

This page carries no code excerpt, and that is a finding rather than an omission. The slice matcher recorded 0 of 8 sections pinned for this page (0 abstention(s), 8 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the eight section keywords. This lane's vocabulary (certification, compliance, audit, logging, monitoring) collides with generic helper and page-component names, and a pinned generic would have given the page the shape of a verified article with none of the substance. Every section above is therefore written from the public sources listed.

Product claims were read from the product source at the revision recorded in this project's pipeline_results.json, read-only, and the plan figures were re-verified against the live pricing page on 2026-09-30. Each heading's keyword comes from this project's own paid measurement run. No code, batch fingerprints, auction data or internal hosts are transcribed.