SmartGate

AI Agent Governance: Who Can Call What, Where Limits Live

AI agent governance is the control plane around an agent, not a document about one. Four decisions do the work: which credential a call carries and what it is scoped to, which layer enforces the rate limit and the token budget, where a policy is evaluated once instead of restated in every prompt, and what the record can prove afterwards.

Short answer: AI agent governance is the control plane around an agent, not a document about one. Four decisions do the work: which credential a call carries and what it is scoped to, which layer enforces the rate limit and the token budget, where a policy is evaluated once instead of restated in every prompt, and what the record can prove afterwards. A rule that lives only in the system prompt is editable by the thing it constrains; a rule in a layer the call must pass through holds.

Key takeaways

  • Governance is enforcement, not paperwork. A rule no component checks is a preference.
  • Four questions, four owners: identity, scope, quota and evidence — and the most permissive layer is the one that wins.
  • Guardrails in an instruction block are requests. The boundary belongs in the credential and the tool allow-list.
  • A budget belongs where the identity is known. A provider invoice cannot say which agent spent the money.
  • Over-reach shows up in the record first: refusals per credential, new tool names, spend velocity.
  • Do this next: list every credential your agents hold, with the tools it may reach and its owner.

ai agent governance is a control plane, not a policy document

The phrase covers two different things, and the difference decides whether anything changes. One is the document: the acceptable-use statement, the risk register, the quarterly review. The other is the control plane: the components that answer a request at the moment it arrives. A document is judged by its coverage; a control plane by whether it can refuse.

Four questions are the whole of it, and each needs exactly one owner.

Question What it decides The place that can answer it What breaks when nothing does
Identity Which principal the call is billed to and recorded under The credential the call carries, resolved by the layer that accepts it One shared key for the fleet; nothing is attributable
Scope Which tools and paths that principal may reach An allow-list evaluated outside the model A read-only agent holds the writer's credential
Quota How much of the ceiling is already spent A counter kept against the tenant, shared between instances Two replicas each spend the whole limit
Evidence What proves a call happened, under whose identity One row per call, written by the layer that enforced Incidents get reconstructed from provider invoices

Name the owner rather than the rule, because a question answered in two places is answered by the more permissive one: a tool's own scope check is the one an operator forgets to update, and an application counter is the one that resets. Pick the layer, make it the only place that decides, and treat every other mention as documentation.

Two properties make a control plane cheap to build. The enforcement point must be outside the thing it constrains — the reasoning loop that wants to finish a task is the loop that would raise its own limit — and it must be on the path, not beside it: a check that runs after the call produces a report about spend rather than a control on it.

How the enforcing layer handles a hostile request is the cluster's centre page: how a policy is parsed and clamped, how a key is validated, and how a payload is masked before it reaches a log are the mechanics of secure prompt handling in AI applications. This page works one level up, on where those decisions are placed and who owns them. Demand for the phrase is modest — 260 searches a month in our own measurement — because the practice is well ahead of the vocabulary: most teams arrive here from "our agents have too much access", not from a procurement line item.

agent guardrails: the boundary has to live outside the model

For an agent, a guardrail constrains what it may do: which tools it may call, which arguments are acceptable, when it must stop, and which decisions a human has to make. What decides whether it works is not how it is written but where it runs.

In the instruction. A line in the system prompt, or a warning in a tool description, is the cheapest guardrail and the weakest: a request to a probabilistic system that reads tool output, documents and user text through the same channel. The OWASP Top 10 for LLM Applications treats prompt injection as a first-class failure mode because instructions and data are not separable at the model's boundary (LLM01). "Do not write to production" is a hope with good grammar.

In the application code around the tool. A tool that validates its arguments, refuses writes on a read path, or stages a change behind a dry run is a real guardrail — deterministic, testable in ordinary unit tests, failing closed. Its blind spot is the caller: the tool sees what it was asked to do, not who was allowed to ask.

In the credential and the layer in front of the tool. What a key is scoped to, and what a layer that knows the key will refuse, is the strongest placement, because it does not depend on the model cooperating or on the argument arriving in the expected shape. OWASP names this pattern excessive agency (LLM06): give the agent the minimum set of tools for the task, prefer an operation that is reversible, and require approval for the ones that are not.

Containing a call the scope does allow is the runtime half of the same problem: which MCP servers an agent trusts, how a tool definition is pinned so an approved tool cannot be swapped underneath it, and how untrusted content is marked before the model reads it are worked through on AI agent security.

That last requirement is where most agent guardrails are actually implemented. The Model Context Protocol lets a server describe a tool's risk on the wire — annotations such as whether a call is destructive or idempotent (MCP specification) — so a client can route it to a human. Treat the annotation as an input to a policy, never as the policy: it is the tool author's claim about their own tool.

Keeping those annotations, and the rest of the published contract, true to the code that serves the call is its own discipline, covered on API documentation best practices for AI tools.

Two properties separate a guardrail from a comment: a stop condition — a maximum step count, a token ceiling or a wall clock, because every runaway agent story is a missing stop condition — and tests that try to bypass it.

llm guardrails: filters reduce risk, they are not the enforcement point

The other common meaning of the word is the model-facing filter: a classifier or validator that runs before the prompt reaches the model, after the completion comes back, or at each turn of a dialogue.

The shapes are consistent across the market. Input rails classify the incoming request — jailbreaks, prompt injection, disallowed topics, personal data that should not leave the building. Output rails classify what the model produced — toxicity, another tenant's data, or claims unsupported by the retrieved context. Dialogue rails constrain the shape of a conversation, which is how a support agent is stopped from promising a refund. NVIDIA's NeMo Guardrails expresses these as flows (NeMo Guardrails), Guardrails AI ships them as validators (guardrails-ai), and independent comparison work has been blunt about the spread in quality across platforms that claim the capability (Unit 42).

Two properties decide how far a filter may be trusted in a governance design.

A filter is probabilistic and sits on the latency path. It can be evaded by a formulation it has not seen, and it produces false positives that break honest requests. That is an acceptable trade for reducing the blast radius of the language layer; it is not an acceptable basis for a control you promise an auditor, because the answer to "can this be bypassed" is yes.

A filter is a decision, so it belongs in the record. If a rail refused a request, the refusal is a governance event: which rail fired, on which call, for which tenant. A filter whose refusals are not counted cannot be told apart from one that was never wired in — treat a refusal count of zero as a defect report, and make the filter fail closed when the classifier errors out.

The division of labour is defence in depth with different jobs. The model-facing rails reduce how often bad language reaches a tool; the deterministic controls — credential scope, tool allow-list, quota, approval for irreversible operations — decide what happens when it does. When a check can be expressed as a rule — an allow-list, a schema, a maximum amount — the rule is cheaper, faster and auditable. The NIST generative-AI profile frames the same discipline as measuring and documenting risk rather than asserting its absence (NIST AI RMF).

token budget: where the ceiling lands, and the ledger behind it

A budget is only real when the same identity pays for every call. Two hosts cannot share a ceiling if the cap is per machine: each host enforces its own limit happily and the org-wide number is discovered on the invoice.

Three candidate places exist for the counter, and one survives contact with a deployment. A provider-side spend cap is account-wide and cannot attribute spend to an agent, so it is an emergency brake rather than a budget. An in-process counter is the easiest to write and the first thing to break: it resets on deploy, and behind two replicas each instance spends the whole limit. A counter in shared state, keyed on the tenant, is the only shape that stays correct across instances and restarts — which is why the enforcement point is usually the same layer that holds the credential. How a credential is scoped to a team, and why the team on a call comes from the key rather than the request body, is worked through on MCP OAuth and authorization.

Enforcement refuses a call; governance has to be able to reconstruct one. That is a different job, and in a gateway it is a read model over two sources that arrive separately — the per-day usage rows, and the cumulative budget actuals the billing side keeps per date:

# lib/dashboard/reports-trend-metrics.ts — source lines 94–122 (buildTokenBudgetSeries)
function buildTokenBudgetSeries(
  dailyTokens: ReportsTrendDataSlice["dailyTokens"],
  budgetActual: ReportsTrendDataSlice["budgetActual"],
) {
  const byDate = new Map<
    string,
    { date: string; input: number; output: number; cumulative: number }
  >();
  for (const row of dailyTokens) {
    byDate.set(row.date, {
      date: row.date,
      input: row.input,
      output: row.output,
      cumulative: 0,
    });
  }
  for (const row of budgetActual) {
    const prev = byDate.get(row.date);
    if (prev) prev.cumulative = row.cumulative;
    else
      byDate.set(row.date, {
        date: row.date,
        input: 0,
        output: 0,
        cumulative: row.cumulative,
      });
  }
  return [...byDate.values()].sort((a, b) => a.date.localeCompare(b.date));
}

The decisions here concern what happens when the two sources disagree about which dates exist. A date may have usage and no budget row, or a budget row and no usage — and the function keeps every date either source mentions, filling the missing side with explicit zeroes rather than dropping the date. That is what stops "no row" from being silently read as "no spend": had the merge skipped the unmatched dates, a month would look quieter than it was, which is the one error a budget report must not make. The final sort is not cosmetic: both loops insert into a map, whose iteration order is insertion order rather than chronology, so without it a chart would draw the merge order and call it time.

The honest boundary is that a series like this is the read side. It runs after the call and answers "where did the month go"; the refusal happens earlier, at the counter, and a dashboard that renders this series is not a control. Building the view first and the check second produces a very clear picture of an overspend you could not prevent.

The read model a budget review actually opens — a call priced against a team, a window and a monthly budget key — is worked through with the shipped functions on AI financial analysis for agents.

Our plans give the ceiling its shape: monthly token caps of 2M, 20M, 100M and 200M+, per-key MCP request rates of 120, 300, 600 and 1200 a minute, and audit retention of 7, 30, 90 or 180 days. Those are operational limits rather than features, and the pricing page is the authoritative table — check it before you commit to a number in a review.

agent observability is how the control plane is verified

A control plane you cannot inspect is a belief system. The rows a gateway or a tracing layer writes are the only way to answer the question governance cares about: did the rules we configured run the way we think they did. What belongs in a row — the fields that make one tool call attributable — is the subject of LLM observability; what follows is what a governor reads out of them.

  1. Does every call carry an identity? A record with no credential behind it means a path that never reached the enforcement point. That is not a logging gap; it is a route around the control plane.
  2. Which credentials were used this week, and which were not? A live key with no calls and no expiry is a liability waiting for a leak; a key that stopped being used usually means the workflow moved somewhere nobody is watching.
  3. Does the tool set per credential still match the job? Scope creep arrives as a new tool name in an existing key's record — sometimes because an agent discovered it, sometimes because a tool description changed and the model started choosing differently.
  4. How many calls does one task need? Calls per task climbing over a baseline mean retries or a loop, and a loop is billed like work: the same intent is paid for several times.
  5. What is the refusal rate per credential? A 403 says the role is wrong for the operation; a 429 says the code is faster than the limit.

Two properties make those readings trustworthy: the tenant filter belongs in the query rather than in a dashboard's default view — if the team predicate is optional, a tenant boundary is a UI setting — and an agent's calls must be distinguishable from a human's. An agent calling with a human's credential looks exactly like a person in the record, and no dashboard work will separate them afterwards.

ai agent monitoring: four signals and the baseline they need

Monitoring turns the record into something that wakes somebody up: a recurring check, a threshold, and a route for the alert. Four signals earn their place, because each is cheap to compute from rows a gateway already writes and each maps to a different failure.

Signal Computed as What it usually means Where it routes
Refusal spike per credential 403 and 429 counts per key per hour A loop, a mis-scoped role, or someone probing a leaked key Read the last twenty calls for that key
Tool-surface drift New tool names per key in 24 hours Scope creep, or a tool description that changed model behaviour Compare the key's intended tool list with its record
Spend velocity per team Tokens this hour against the trailing 7-day median The work changed shape, or an agent is retrying in a loop Read the task with the most calls, not the most tokens
Credential hygiene Age, missing expiry, calls in the last 30 days Keys nobody owns anymore Reissue with an expiry, or revoke and watch what breaks

Monitor refusals rather than successes, because a working limit is invisible in the happy path. A key that quietly hits its per-key request window several times an hour is a key whose workflow does not fit its plan — a capacity problem presenting as a governance problem.

Two failure modes are worth designing against: alert fatigue, because a refusal is often a limit doing its job, so threshold the rate rather than counting refusals; and an alert with no owner, because a signal routed to a channel nobody reads is a log line with extra steps.

One signal is easy to forget because it is not about agents at all: monitor the control plane itself — who raised a cap, who added a key, who edited the policy, and from which role. Those changes make every other reading mean something different, and if they are not recorded the record describes a system whose configuration cannot be dated. In our stack that is a property of the enforcement point rather than of a separate tool, because the policy edit and the call it governs are enforced and recorded by the same layer — the arrangement described on MCP gateway.

ai audit trail: attribution, cadence, and what the record cannot prove

A governance trail has one narrow job: pulling one credential's activity over one window and reconstructing what it did. Everything else the phrase suggests — a retention schedule, an export format, a statement for an auditor — is the evidence discipline owned by the compliance side of this cluster; the governance use is narrower.

That evidence discipline — the fields an audit row must carry, how a retention period is chosen and how the record is exported to an auditor — is set out on AI compliance.

Attribution is the property that makes the rest possible. A row is attributable when it names the credential that authenticated the call, and a task is attributable when its calls share a correlation key: without one, a five-step agent run is indistinguishable from five unrelated users of the same tool. That key also lets a reviewer read a run in order.

The cadence that works is boring. Daily, automatically: the four signals above. Weekly, by a human who owns an agent: pull one credential's week, replay the sequence, and ask three questions — was this tool list intended for this workflow, was this amount of spend expected, and could we explain this run to a customer who asked. Quarterly, by whoever owns the risk: does every live credential still have an owner, an expiry and a reason. The weekly review is the one that finds things, because it reads intent against records.

Two boundaries keep this honest. The first is that the trail proves the credential was valid, not that the call was intended: an agent doing the wrong thing with the right key produces ordinary-looking rows. The trail is what lets you reconstruct a sequence, attribute it and stop it; it is not a detector of purpose, and treating it as one is how a team concludes after an incident that the logs looked normal. The second is the window: retention is an entitlement rather than a preference — 7, 30, 90 or 180 days depending on the plan — so it also bounds how far back a review can look, which makes it a decision about worst-case discovery time rather than average storage cost.

There is a third boundary that is really an argument for one path: the trail describes exactly the traffic that went through the thing that wrote it. A host that talks straight to a model provider is absent from the record and invisible to every check on this page.

Where SmartGate fits

SmartGate is an MCP gateway, which in this page's vocabulary means it is an enforcement point and the record keeper for the same calls. Every tool call through it resolves to a key, a team and a plan, is written as an audit row, and is measured against the limits that key and team carry — so the readings above are queries over the traffic you routed through it rather than instrumentation you have to add.

The same server is what an individual pastes into their own assistant for a research workflow, described from the user's side on SmartGate MCP for research and decisions.

For a control plane that gives four usable properties: identity comes from the credential rather than the request, the effective policy is whichever of the plan and the team policy is tighter, the counters live in shared state so a limit holds across instances, and each call writes the row a review reads. The plan sets the operational numbers rather than the feature list — monthly token caps of 2M, 20M, 100M and 200M+, MCP requests per minute per key of 120, 300, 600 and 1200, audit retention of 7, 30, 90 or 180 days, and 2, 10, 30 or unlimited keys per team. Compare those against your worst-case discovery window and your busiest agent's tool list before you commit to a date; the pricing page is the authoritative table.

Two things it deliberately is not: it does not classify prompts or judge answers, which is the filter layer above and a separate purchase in most stacks; and it does not decide what your agents may do, because the tool list, the scope and the review cadence are policy.

Frequently Asked Questions

Limitations

This page describes a control plane, not a compliance program. It does not tell you what your regulator requires, what to retain, or how to produce evidence — those questions belong to the evidence side of this cluster, and the answers depend on your jurisdiction rather than on your architecture.

Nothing here is a benchmark of guardrail products. The two guardrail sections are written from public standards and vendor documentation because the slice matcher could not pin a unique symbol for either phrase in the repository it scanned, which the Method note records; no product is ranked, and the claim that filter quality varies widely is the cited comparison work's, not ours.

The one code excerpt covers a read model — the date-keyed series behind a budget view — and says nothing about the enforcement path: where the counter lives and how a refusal is shaped are not quoted here.

The monitoring signals are leads, not thresholds: every number in that table has to be calibrated against your own traffic before it alerts on anything. The plan figures above were re-verified against the live pricing page on 2026-09-30 and are operational limits rather than a feature comparison, so read the current table instead of this page.

Sources

Method note

The code in this article is not transcribed. The fenced block was cut out of the slice body returned by the slice API and re-asserted byte-for-byte as a substring of that body before publication, and the first line inside the fence records the file and the exact source lines. The symbol was pinned by whole-name containment (rule A level 3) and confirmed by the slot-proof endpoint; the window is the whole slice body.

Only one of this page's seven sections pinned a code excerpt, and the two guardrail sections are quoted from nowhere at all on purpose. Rule A matched the generic fixture name guard inside a test file for both agent guardrails and llm guardrails, slot-proof returned no asset for that match, and the pin stage recorded an abstention rather than ship a generic name under the article's two most standard-bearing headings. Those sections are written from the public standards, specification and vendor documentation listed above, with no code quoted and no claim about how any product implements a rail. The remaining four sections carry no excerpt either, because rule A found no unique symbol behind their vocabulary.

Product claims were read from the product source at the revision recorded in this project's pipeline_results.json, read-only, and the plan figures were re-verified against the live pricing page on 2026-09-30. The section keyword behind each heading comes from this project's own paid measurement run, not from a third-party tool. No code beyond the excerpt above, batch fingerprints, auction data or internal hosts are transcribed.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)
1 token budget buildTokenBudgetSeries lib/dashboard/reports-trend-metrics.ts 94–122 rule A L3 → slot-proof e038d17dd099

Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before publication. 1 of 7 sections pinned, 2 abstentions, 4 misses.