SmartGateSmartGate

Agent Identity and Authorization: NHI Design for AI Agents

An agent identity is the non-human identity an autonomous workload uses to authenticate itself, distinct from the human account that authorised it. Because an agent acts on its own between check-ins, it needs a credential scoped to one job, short-lived enough to rotate without a person, and recorded on every action so the audit trail can answer who did what.

Short answer: An agent identity is the non-human identity an autonomous workload uses to authenticate itself, distinct from the human account that authorised it. Because an agent acts on its own between check-ins, it needs a credential scoped to one job, short-lived enough to rotate without a person, and recorded on every action so the audit trail can answer who did what. Agent identity is those three decisions taken together — how the workload authenticates, what its credential is allowed to authorise, and which identity the record attributes each action to — and the multi-agent case adds a fourth: how trust and narrowed authority pass from one agent to the next.

Key takeaways

  • An agent identity is a workload identity, not a person's account. Reusing a human login for an agent makes the credential impossible to disable alone, over-permissioned by default and useless for attribution.
  • Authentication and authorization are separate questions. A valid token says which workload is calling; a scope and a policy decide what that workload may do, and a valid token is never a permission by itself.
  • The lifecycle is where deployments fail. An identity with an owner, an expiry and one place to revoke is manageable; a static key pasted into five agents is five leaks and one rotation.
  • Delegation must be explicit. When one agent calls another, exchange a token and record the acting party rather than passing the original credential down the chain.
  • Do this next: list every agent that can call a tool, give each one its own scope, and read a single audit row to confirm which identity produced the action before trusting the fleet.

What agent identity is, and why it is not a user account

An agent identity is the account a software workload authenticates with when it acts on its own. The definition is narrow on purpose: it is not the human who configured the agent, not the team that pays for it, and not the model behind it. It is the principal that surrounding systems check when a tool call arrives, and the name that ends up in the log line afterwards. Once an agent can call tools, reach private data and write outward, "which identity made this call" stops being a logging detail and becomes the control everything else — limits, scopes, attribution — is attached to. It is also the credential an attacker most wants, which is why the same context that makes a prompt injection worth planting is what makes a stolen agent identity worth using.

A user account is a poor fit because the two identities fail differently. A human credential is carried by a person who can be challenged, phished and prompted for a second factor; it is expected to roam across devices and to be reset when a laptop is lost. An agent runs unattended, often in a container or a scheduled job, and may execute thousands of calls between two human decisions. Reusing a personal account for that workload produces the failure every identity team recognises: the account cannot be disabled without disabling the person, its permissions are the union of everything the person ever needed, and the audit log cannot separate a human's click from a batch job's run.

The security industry calls this class of identity non-human, and the agent is its newest and fastest-growing member. That is why the next section treats the general category before narrowing back to agents — the agent inherits the category's problems and adds properties of its own.

Non human identity: the workload account, one layer down

Non human identity (NHI) is the umbrella term for every account a machine rather than a person authenticates with: service accounts, API keys and client secrets, service principals, workload and managed identities, certificates, and the tokens automation holds on their behalf. Microsoft's security guidance groups them exactly that way — identities that software, services and automation use to access resources — and it points at the operational fact that gives the category its urgency: non-human identities typically outnumber human ones by a wide margin in a modern estate, while the controls built for human accounts assume a person is present to respond (Microsoft Security 101). IBM's treatment makes the same point from the identity-management side — an NHI is any identity not tied to a human, and it needs its own lifecycle, ownership and review because nobody will ever log in to notice that it should have been retired (IBM).

That asymmetry is the design constraint. A human account gets attention by default: someone uses it, notices a prompt, reports a lockout. An NHI gets attention only if the system is built to give it — an owner recorded at creation, an expiry that forces a decision, a review that surfaces the ones nobody uses. An agent is a non-human identity with two extra properties that sharpen the general problem. It is created and discarded far faster than a service account, and it carries authority a human delegated to it, so its credential is a place where a person's permissions can outlive that person's oversight. The model and dependency supply chain is the sibling topic that covers where the agent's weights and libraries come from, which is a different question from who the agent is.

The category exists because the controls were built for people. A password reset, a second-factor prompt and a session timeout are all shaped around a human present at a keyboard, and none of them transfers cleanly to a workload that has no keyboard to sit at. That is why NHI management grew its own vocabulary — secrets management, machine identity, workload identity — and why the questions worth asking about an agent are lifecycle questions rather than authentication questions: who owns this identity, when does it expire, and which single action revokes it.

Agent authentication: proving which workload is calling

Authentication answers a single question — is this really the workload it claims to be — and the menu of mechanisms is now much wider than the shared API key most agents still use. The OAuth 2.0 client-credentials grant lets a workload obtain a token with its own client identifier and secret, with no user present (RFC 6749). The Model Context Protocol's authorization specification builds on that world: HTTP-transport servers are expected to behave as OAuth 2.1 resource servers and to publish protected-resource metadata so a client can discover the authorization server, and its security guidance forbids passing a caller's token through and accepting tokens in a query string (MCP authorization, MCP security best practices). That protocol surface is worked through in full on the MCP authorization page; this section treats it as one option and asks which option fits an agent.

Three upgrades matter more than the choice of flow. First, bind the credential to the workload with a key it holds rather than one it merely carries: mutual-TLS-bound tokens and DPoP make a stolen token useless without the corresponding private key (RFC 8705, RFC 9449). Second, prefer federation over a stored secret: Google Cloud and Microsoft Entra both let a workload exchange an external, attested identity for a token with no long-lived secret on disk, and AWS issues temporary credentials through STS instead of an access key that lives forever (Google Cloud workload identity federation, Microsoft Entra workload identity federation, AWS IAM roles). Third, treat the identity as the thing that rotates, not just the secret: SPIFFE/SPIRE issues each workload a short-lived cryptographic identity that is renewed automatically, so there is no static artifact to exfiltrate from a config file (SPIFFE overview). The weakest option — one long-lived key pasted into every agent's configuration — is also the one that makes attribution impossible, because every call arrives under the same name.

Agent authorization: from a valid credential to a permitted action

Authentication says who is calling; authorization decides what that identity may do, and conflating them is the most common mistake in agent platforms. A valid token is not a permission. OAuth scopes are coarse — read and write on a service, not permission for one record — and the resource indicator narrows a token to an intended audience so it cannot be redeemed somewhere it was never meant to go (RFC 8707). Least privilege is the NIST discipline of granting the minimum needed and reviewing it: the Zero Trust Architecture removes implicit trust from network location and asks for an authorisation decision per request, and NIST SP 800-53 states least privilege as an access-control control rather than a best-effort (NIST SP 800-207, NIST SP 800-53). For an agent the practical translation is an inventory: list every tool it can call and every destination it can write to, then delete the entries the task does not need.

The failure mode to design against is the confused deputy, where a more privileged service is tricked into using its own authority on behalf of a less-privileged caller. MCP's security guidance names it directly and pairs it with the token-passthrough anti-pattern, in which a server forwards the client's token downstream instead of minting a token of its own with the audience it actually needs. The rule that follows is simple to state: every distinct authority an agent exercises should be a distinct token with the narrowest scope and audience that still lets the step work. A model can be talked into an action by content it reads — the mechanism behind jailbreaking a model's own policy is adjacent — but it cannot mint authority it was never issued, which is why the authorisation layer has to live outside the prompt.

Agent identity management: issuance, rotation and revocation

Identity management is the lifecycle around the credential, and it is where most agent deployments are weakest, because the fast path — create a key, paste it into a config, move on — has no lifecycle at all. NIST SP 800-53 treats account management as a process with defined stages: authorise the account, establish it with the right privileges, monitor it, and remove or disable it when the need ends (NIST SP 800-53). The agent version of those stages is short and concrete. Issuance should attach an owner and a purpose, because an account with neither becomes an orphan that no review will ever flag. Lifetime should be measured in hours or days, not years, so a leaked credential expires before it can be exploited at leisure. Rotation should be automatic and invisible to the agent, which is what temporary credentials buy: an AWS role issues STS credentials that expire, Google and Entra federation mints short-lived tokens, and SPIFFE renews an identity before the old one lapses (Google Cloud service-account best practices, AWS IAM best practices). Revocation must be reachable in one place — a token-issuing service that can refuse to renew, rather than a search across every host that holds a copy of a secret.

The reason to invest here is the asymmetry between a static secret and a short-lived one. A long-lived key copied into five agents is five places to leak and one place to rotate, while a short-lived token is one place to leak with a lifetime measured in minutes, and bot comments in the vendor glossary pages that dominate this query tend to skip that trade. Orphaned identities are the quiet half of the problem — Microsoft's NHI guidance flags sprawl and unmanaged credentials as the standing risk — and the only reliable cure is an owner and an expiry recorded at creation, so the review has something to act on.

Two failure stories are worth keeping in view because they are opposites. The first is the credential that never expires — a key created for a proof of concept, copied into a staging config and then into production, still valid years later and readable by anyone with repository access. The second is the credential that expires in the middle of a run — a lifetime chosen without regard for the longest task the agent performs, so a long job fails halfway and the tempting fix is to lengthen the lifetime until rotation is effectively gone. The design answer to both is the same: a lifetime tied to the task, a renewal the agent can perform without a human, and an alert when a credential has not been renewed on schedule.

AI agent access control: scopes, least privilege and boundaries

Access control for an agent is the set of rules that decide which call proceeds, and the useful move is to build it as a boundary outside the model rather than a sentence inside the prompt. Three layers carry most of the weight. The first is scope granularity: coarse scopes like read and write are too blunt once an agent holds dozens of tools, so the practical unit is a resource plus an action, and the token an agent holds should carry only the units its current task needs. The second is a read/write split, because an agent that can only read cannot be made to exfiltrate by an instruction it reads. The third is deny-by-default policy for mutating actions, where anything that sends, pays, deletes or publishes requires either a narrow scope or an explicit approval step.

Enforcement has to sit where the call happens, not where the answer is written. That is the difference between describing an agent's limits in its system prompt and enforcing them at the point the tool is invoked: a prompt is a request, and a policy check is a gate. The gateway pattern puts that gate on the path of every call, so the identity's scopes, its rate limit and its budget all apply at the same instant, and the same component writes the record of what happened. It is worth separating this from containment: where untrusted code runs is isolated by agent sandboxing, a different control with a different failure mode; access control decides what the identity may do wherever the code happens to execute, and the two are complements rather than substitutes.

Agent access control when agents call other agents

Multi-agent systems turn identity from a property of one workload into a property of a chain, and the chain is where authority usually leaks. When a supervisor agent dispatches a task to a worker, three designs are possible and only one of them is auditable. The first passes the supervisor's credential down; the worker then acts as the supervisor, with the supervisor's full authority, and the log shows only the originating identity. The second gives each agent its own identity and lets the worker use it against a shared resource; that is better, but the worker still does not record that it is acting on the supervisor's behalf. The third uses token exchange: the worker presents its own token plus the incoming token and receives a new token whose actor claim records that it is cooperating on behalf of the original principal, so delegation is explicit rather than assumed (RFC 8693). Microsoft's on-behalf-of flow is the same idea in an Entra estate — a middle tier exchanges a user's token for a token it can use downstream (Microsoft Entra on-behalf-of flow).

The rule that makes the chain safe is that authority narrows at each hop and never widens. A worker should receive a token with a subset of the supervisor's scopes and the exact audience of the service it will call, so a compromised worker cannot use the token anywhere the task did not require. Trust between agents also wants a boundary, which SPIFFE/SPIRE's notion of a trust domain supplies: each workload in a domain gets an identity its peers can verify. The anti-pattern is the shared team token every agent in a graph can read; it makes the fleet easy to wire and impossible to audit, and it is also the easiest target for instructions that arrive inside fetched content, because one poisoned input is then enough to move the whole crew. The MCP-specific transport and delegation details stay on the authorization page linked above.

Audit attribution: whose identity is on the action

Attribution is the question an incident review actually asks: not "what happened" but "whose identity did this". Getting it right means the record has to carry the acting identity for the specific call, not the human who owns the fleet or the team that pays the bill. A token's subject identifies the principal; a delegation chain recorded with it — the actor claim an exchanged token carries — identifies who is acting on whose behalf. A single log line can then distinguish a human's own action from an automated one performed for them, and one identifier threaded through the calls of a single task lets a reviewer reconstruct the sequence without guessing which tool call belongs to which request.

Two design choices make attribution durable. First, log the credential identity and its scope on the action, so the record says which key, not merely which team. Second, log the delegation, because an agent that acts for a user is a different event from the user acting alone — the confused-deputy case the MCP guidance warns about is precisely a log that records the wrong party. This is the identity half of the record; the policy half — how long rows are retained, what a compliance export must contain, which claims may not be made — belongs to the wider LLM security control plane and its compliance siblings rather than being repeated here.

How identity travels with a call, read from our own context

"Attribute the action to an identity" sounds like one field. Ours is a set of request-scoped values, read on 2026-10-07 from backend/smartgate/core/request_context.py and core/audit_enrichment.py.

  • Identity is bound once at the edge, then read. Team id, route, agent platform, transport, key id, source address, correlation id, trace id, event kind and even the segment index of a chunked call are bound into a request-scoped context as the call enters. Every later consumer — the audit hook, a module, the refusal path — reads the same values instead of re-deriving them. Re-derivation is where attribution drifts: two components that each guess the caller eventually disagree.
  • The agent platform is its own field, not a guess. Which framework a call came from is bound explicitly rather than inferred from a user-agent string, so "which agent did this" is answerable without a parser and a regex you will regret in six months.
  • Rotation and provenance are separate fields too. The key id is bound alongside the team, so a record can say which credential acted — the difference between "the team did this" and "this key did this", which is the first question in any incident review.
  • Trace identity flows in from outside when it exists. The context takes the correlation id from the incoming headers (with a documented fallback) and the trace id from the session map or the client's own header — the design rule being that an identifier created at the edge is better than one created later, because only the first can be shared with the caller.

Where SmartGate fits

SmartGate does not try to be the identity provider for your fleet, and this page will not claim otherwise. What it provides is the enforcement and attribution point for the calls an agent makes through the gateway: a single authenticated endpoint where a per-key identity, its scope, its rate limit and its audit row meet at the moment of the call. That is useful precisely because the identity decisions above need somewhere to land — a key that can be scoped and revoked is an agent identity you can manage, and a call written as an audit row with its key is attribution you can read back. Seven tools are exposed through the gateway — smart_fetch, smart_search, smart_context_gate, smart_dedup, smart_budget_guard, smart_memory and smart_pipe — and every call is counted against the identity that made it. Call control is documented on token control, the record on audit and compliance, and the plan limits move with the tier: monthly token caps of 2M, 20M, 100M and 200M+, requests per minute per key of 120, 300, 600 and 1200, audit-log retention of 7, 30, 90 or 180 days, and 2, 10, 30 or unlimited keys per team. The pricing page is the authoritative table.

How to get started

  1. Inventory the agents. List every workload that can call a tool or reach a resource on its own, and name the human who owns each one. An identity you cannot list is one you cannot revoke.
  2. Give each agent its own identity. Split shared service accounts and team keys so one credential maps to one workload; the moment two agents share a key, attribution is gone.
  3. Replace static secrets with short-lived credentials. Move to federation or temporary credentials where the platform supports it; where a secret must stay, give it an expiry and an automatic rotation path.
  4. Scope each identity to its task. Start with a resource-and-action scope, remove the tools the agent does not use, and split read from write so a read-only agent cannot exfiltrate.
  5. Record the acting identity on every call. Log the key and its scope, and thread one identifier through a task's calls so a review can reconstruct the run.
  6. Enforce at the gate. If calls run through a gateway, start free and confirm that a revoked key stops the next call before you trust the limits.

Frequently Asked Questions

Is an agent identity just a service account with a new name?

Mostly yes at the mechanism level, and the distinction is about dynamism rather than kind. Agent identities share the properties of service accounts — no human, programmatic use, a credential that must be issued and revoked — but they are created and replaced far more often, and they usually act with authority delegated by a person for a specific task. That delegation is what makes attribution and narrowing scope the central problems rather than an afterthought.

Can one agent reuse a human user's account?

It works until you need to answer a question about it, and then it fails on every axis. A human account cannot be disabled for the workload without disabling the person, its permissions are the union of that person's access rather than the task's, and the log cannot separate a human action from an automated one. For a short experiment it may be acceptable; for anything that runs unattended, it is the mistake the rest of this page exists to prevent.

Why are static API keys still so common, and what replaces them?

They are common because they are the fastest thing to wire and the hardest to notice breaking. What replaces them is a credential the platform issues for a short interval and can refuse to renew: temporary credentials from a cloud role, a token obtained through workload identity federation, or a SPIFFE identity renewed automatically. The migration is worth doing because a static secret copied into several agents is several places to leak and one place to rotate, while a short-lived token narrows the window itself.

How should one agent authorise a call made by another agent?

By exchanging a token rather than sharing one. The acting agent presents its own identity and the incoming token, receives a new token that records the original principal on whose behalf it is acting, and uses that token with a scope narrowed to the specific service it is calling. Passing the original credential down instead collapses the chain into one identity and makes the delegation invisible in the audit log.

Does enforcing identity at a gateway replace least privilege in the agent's own code?

No, and treating it that way is a common overreach. The gateway enforces the limits attached to an identity — scope, rate, budget — at the moment of the call, which is where a central check is strongest. It does not decide which tools the agent should have been given or whether a particular task needed a write scope; that is a design decision made when the identity is created. The two layers answer different questions and should both be present.

Limitations

This page is a design framework, not a product comparison, and it deliberately names no identity provider as the right answer. The mechanisms it describes — OAuth-based authentication, token exchange, workload identity federation, SPIFFE-style cryptographic identity — are options with different operational costs, and which one fits depends on the platforms an agent already runs on rather than on a general preference.

It also does not claim that a strong identity model removes the other risks an agent carries. A well-scoped identity narrows what a compromised agent can do; it does not stop the agent from being manipulated into doing something within that scope, which is why access control is a complement to containment and to input-side defences rather than a replacement for them. The external descriptions above are each source's own published wording, read at the linked pages, and the demand figures are this project's own measurement rather than a third-party estimate.

Sources

Method note

This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned none of this page's seven sections, returning seven no-slice verdicts and no abstentions: rule A found no unique symbol in the scanned repository for any section keyword, and the remote candidate fallback returned generic helpers — a request-context field, an MCP header builder, a cron guard, memory accessors — that are collisions rather than section-specific evidence about agent identity. A pinned generic would have given the page the shape of a verified article with none of the substance, so every section but one above is written from external, linkable sources, the house rule for an unpinned section.

Product facts were read read-only from the product source at the revision the slice run recorded in this project's pipeline_results.json, and the plan figures were re-checked against the live pricing page on 2026-10-04.

The slice run for this page recorded 0 of 7 sections pinned, 0 abstention(s) and 7 no-slice verdict(s); BLOCKS is empty because the matcher found no unique symbol for any section rather than section-specific evidence, as the Method note above explains.