LLM Security for Platform Teams: Keys, Logs, and Controls
LLM security for a platform team is mostly a control-plane problem, not a model problem. What you can actually enforce lives outside the model: which key made a call, what it is allowed to spend, what the log may contain, and what the gateway does when a caller crosses its limit. This page is the operations checklist for those controls, from key custody to dependency hygiene.
Short answer: LLM security for a platform team is mostly a control-plane problem, not a model problem. What you can actually enforce lives outside the model: which key made a call, what it is allowed to spend, what the log may contain, and what the gateway does when a caller crosses its limit. This page is the operations checklist for those controls, from key custody to dependency hygiene.
Key takeaways
- Most LLM security decisions a platform team owns are about custody and records: who holds a credential, what a log captures, and what a limit refuses.
- A shared key is a shared identity; without per-caller keys you cannot attribute spend, revoke one caller, or hold a budget for one team.
- Log the metadata, not the payload: caller, tool, tokens, latency and outcome, with payload bodies withheld or redacted by default.
- Gateway controls bound blast radius even when nothing detects an attack; rate limits, quotas and an audit trail are enforcement, not observation.
- Start by writing down what may enter a context window and what may be kept afterwards; the policy is the control, and the gateway is where it is applied.
What llm security means for a platform team
Security for a language-model system is usually described from the model outward — jailbreaks, prompt injection, and how an attacker talks a model into misbehaving. Those are real, and the other pages in this cluster take them apart. The platform team owns a different slice, and it is the one that decides how much damage any of those attempts can do. Three questions define it: who is allowed to call the system, what the system records while it runs, and what it refuses when a caller exceeds what it was granted.
The distinction matters because the two kinds of control behave differently. Model-side defenses are probabilistic: a classifier or a hardened model lowers the chance of a bad outcome and never reaches zero. Platform-side controls are deterministic: a revoked key cannot authenticate, a quota that reads zero refuses the next call, and an audit row either exists or it does not. Where the model layer gives you a probability, the control plane gives you a bound, and a serious program runs both.
This page stays on the deterministic side. It does not explain how injection or jailbreak work — those mechanisms belong to the prompt injection threat model and to jailbreak attempts against the model's policy, and re-telling them here would be both repetitive and beside the point. It covers the surfaces a platform or operations engineer can configure and test this week: credentials and key lifecycle, logs and PII, gateway rate limits and quotas, the data boundaries around prompts and outputs, and the supply-chain and patch hygiene that keeps an agent stack current.
ai security starts with who can call the system
Every request into a model stack arrives with some identity, and the first platform decision is what that identity is. If it is a single API key pasted into every service, you have one principal: you cannot attribute spend to a team, you cannot revoke one integration without breaking all of them, and a leaked key is a total compromise of the account. Per-caller keys are the foundation every other control on this page reads from.
The work is a lifecycle rather than a checkbox. Issue a key per environment and per service. Store it in a secrets manager, never in an image, a repository or a build log. Give it the narrowest scope the caller needs, so that a compromise of one service is not a compromise of the account. Rotate on a schedule and on suspicion, and rehearse revocation until it takes minutes — the worst time to discover that nobody knows which service holds key number four is during an incident. Keep a register that maps each key to an owner and a purpose, because attribution is the point.
This is where agent identity and authorization begins; that page works through non-human identity, credential issuance and the audit question of who did what. From the platform side the rule is short: no shared secrets, no long-lived keys checked into code, and a revocation path that is faster than an attacker is. Every call through a gateway is made under a key, and the key is the unit at which cost, limits and records are all measured.
llm security risks that live in your own infrastructure
The risks that a platform review actually finds are rarely exotic, and almost none of them are about the model's weights. They are the ordinary failures of running a system that spends money on every call and reaches data on behalf of users. It is worth naming them as a list, because each one has a control that maps to it.
- Over-broad credentials. A key that can call every tool with no per-team limit turns a single stolen value into an account-wide incident.
- Unbounded consumption. An agent that loops, retries or over-fetches has no natural brake; without a quota, the first sign of trouble is the invoice.
- Logs that over-retain. Copying full prompts and responses into a log store turns a debugging convenience into a data-retention liability.
- Dependency drift. Client libraries, the proxy and model endpoints all move underneath you; a stack pinned to nothing is a stack you cannot patch on purpose.
- Unattributed actions. If the record does not name the key behind a call, you cannot answer the one question every incident review asks.
None of these needs an attacker to begin with; they are misconfigurations, and that is the reason a platform view is worth writing down at all. This is the surface you can inventory, bound and test, and it is where the leverage is. The specific techniques that exploit a system once it is reachable are covered elsewhere in the cluster, so the goal here is to make the infrastructure itself the first line rather than the last.
Logs and PII: the record you keep decides what you can prove
The most consequential decision a platform team makes about LLM security is what the log contains, because that single choice sets both what you can prove after an incident and what you are storing in the meantime. Two logs get conflated and should not be.
An access log records that a call happened: the principal, the route, the tool, the token count, the latency, the outcome and the timestamp. A content log records what was said: the prompt text, the model's output, the tool's return value. The first is safe to keep broadly and cheaply. The second is often regulated data wearing a debugging label.
The default that survives review is metadata by default and content on exception. Record the access row for every call, and keep payload bodies only where a specific purpose requires them, with a retention window and a redaction step attached. If a caller can send personal data, then the prompt and the completion contain it, and storing both without a policy is the leak you will have to disclose. Where content must be kept, strip identifiers on the way in, encrypt it at rest, restrict who can query it, and tie retention to the shortest window that still answers your questions. The plan's audit retention moves across four tiers — 7, 30, 90 and 180 days — so a policy that outlives the plan it runs on is a policy you cannot actually enforce.
Access control over the log matters as much as its contents: who can read it, who can export it, and whether the exports are themselves recorded. The same record that lets you reconstruct a call is the record a customer or a regulator will ask you to produce, which is why audit and compliance is a design surface rather than a setting you switch on afterwards.
Gateway-side controls: token limits, traffic shaping, and audit trails
A gateway is the natural enforcement point because it is the one component every call passes through, and that makes it the only place a limit can be applied consistently. A limit implemented in application code is a suggestion each service may or may not honor; a limit implemented at the gateway is a property of the whole system.
Three controls do most of the work. Rate limiting caps requests per key per minute, so a runaway loop degrades to a known rate instead of saturating an upstream dependency. Quotas cap consumption over a period — tokens per month, or a per-team budget — so that spend has a ceiling the next call cannot cross. Traffic shaping decides how a burst is handled, whether it is rejected, queued or slowed, and whether one noisy caller can affect the others. Each of these is deterministic, and each is testable, which is the property that matters: you can prove a rate limit works, and you cannot prove a classifier will catch what you did not anticipate.
The audit trail is the fourth control and the one that makes the other three accountable. A record written as calls happen lets you connect a spend spike to a key, a key to a team, and a team to a request, without asking anyone to remember what happened last Tuesday. Token control is the surface that carries the limits; the record is the surface that proves they applied.
SmartGate's plan table sets the operational numbers — monthly token caps of 2M, 20M, 100M and 200M+, and requests per minute per key of 120, 300, 600 and 1200 — so the pricing page is the authoritative table to read them against your own volume, and those figures move with the tier rather than staying fixed.
Prompt and output data boundaries
Data boundaries are a platform concern because they decide what leaves your perimeter on every single call, and they are cheap to draw early and expensive to retrofit.
The input boundary is what a caller may place into a context window. If one account can reach both regulated data and untrusted third-party content, you have combined two trust levels in one context, and no platform control can separate them again once they are there. The architectural fix — different accounts or scopes for different data classes — is cheap at design time and painful after a product depends on the combined path. This is the practical reason indirect prompt injection through fetched content matters to a platform team even though the mechanism is not theirs to fix: it is the case that punishes a boundary you did not draw.
The provider boundary is which model endpoints receive which data, and under what terms. Keep a register of destinations, send the minimum a task needs, and prefer the route that does not copy content to a system outside your control.
The output boundary is where completions are stored and who can read them. Treat model output as data your users produced, not as platform metadata, and apply the same retention and access rules you apply to the input. A boundary you cannot describe in one sentence is probably a boundary you do not have.
Deployment and dependency supply chain basics
The last surface is the stack itself: what you run and what it pulls in. An agent stack is typically a client library, a gateway or proxy, one or more model endpoints, and a set of tools, and every one of those is a dependency that can change underneath you.
Four habits keep this manageable. Pin versions for libraries and container images, and upgrade on purpose rather than by drift. Verify what you deploy — signed images, or at least checksummed artifacts — so that a compromised registry does not become a compromised runtime. Keep secrets out of images and build logs, because a credential that reaches a layer or a log has already crossed the boundary you drew above. And run the gateway with the least privilege it needs, placed so that a compromise of one component is not automatically a compromise of all of them.
The same discipline applies to the tools an agent can call, since each tool is a capability the platform should be able to enumerate. Where tool code runs untrusted or third-party logic, agent sandboxing covers the isolation that keeps a failure contained instead of letting it escape; the platform responsibility is to make that isolation the default rather than an option someone has to remember. Supply-chain work is unglamorous, and it is the difference between a stack you can patch deliberately and one you can only hope about.
llm vulnerability and patch hygiene on the platform
The word vulnerability carries two meanings in this context, and keeping them apart prevents a great deal of wasted effort. The first is the ordinary software vulnerability — the CVE in a library, the outdated proxy, the base image with a known flaw. That is standard vulnerability management: keep a bill of materials for the stack, watch the advisories for the components you actually run, and maintain a patch window you can shorten when something is being exploited in the wild. Cadence is the whole game.
The second meaning is the model-level weakness — the class of behaviour in which a model can be steered — and that is not something a patch fully closes. Because it cannot be patched away, the platform's response is not a fix but containment: the same deterministic controls described above, the limits that bound what a misbehaving call can spend, the boundaries that bound what it can reach, and the record that shows what it did. When a new class of weakness is published, the useful platform question is not whether your model is immune but what the worst a single call can do if it is exploited here.
Where that taxonomy lives, and which entries to prioritize first, is the cluster's OWASP Top 10 for LLM Applications walkthrough; this page deliberately does not re-list the categories or their mitigations. The division of labour is simple: that page maps the risks, and this page maps the controls that sit underneath them.
llm security best practices as an operations checklist
A checklist is only useful if every line is testable, so the rows below are ordered so that each one makes the next cheaper to decide, and each names the evidence that shows it is done.
| Area | Control | Evidence it holds |
|---|---|---|
| Identity | One key per service and environment; no shared secrets | A key register with an owner per entry |
| Custody | Keys in a secrets manager, never in images or code | An image and history scan that finds none |
| Revocation | Rotation on a schedule and a rehearsed revoke | A revocation that was timed under pressure |
| Rate limits | Per-key requests per minute enforced | A burst test that gets throttled |
| Quotas | A per-team token or budget ceiling | A test where the next call is refused at zero |
| Logging | Metadata recorded; payloads withheld or redacted | A log query that returns no raw personal data |
| Retention | Windows set per plan and enforced | Deletion happening on schedule |
| Boundaries | Input, provider and output boundaries written down | A one-page boundary register |
| Supply chain | Versions pinned and images verified | A bill of materials and a patch log |
| Attribution | Every action names the key that made it | An incident reconstructed from the record |
Two rows deserve emphasis because they are the ones teams most often skip. Revocation is a control only if it has been timed under pressure; an untested revocation path is a plan, not a control. And attribution is what turns the whole table from prevention into detection: the controls stop the calls you anticipated, and the record is what catches the ones you did not.
llm security testing and ai security testing of your own controls
Testing an LLM deployment splits into two disciplines, and a platform team owns only one of them. The discipline it does not own is adversarial testing — probing a model to make it misbehave — which is a specialist method with its own tooling and its own place in this cluster. The discipline it does own is control testing: proving that the deterministic controls you configured actually behave the way the policy says they do.
Control testing is fast and unglamorous, and it converts a policy document into evidence. Test the quota by exhausting it and confirming the next call is refused. Test the rate limit by bursting past it and confirming throttling rather than saturation. Test revocation by disabling a key and confirming the call fails closed, not open. Test redaction by sending a known marker through a call and grepping the log to confirm the marker is absent. Test attribution by making a call from a known key and finding it in the record. Each of these takes minutes, and each produces a fact you could hand an auditor.
The two disciplines answer different questions. Adversarial testing asks whether the model can be steered; control testing asks whether the guards around it hold when it is. A program with only the first has a threat model and no proof of enforcement; a program with only the second is well instrumented and unaware of what it is defending against. The mature answer runs both on a cadence, and keeps the control tests in the same suite as the rest of the infrastructure checks, so that a change to a limit or a key policy is caught by a failing test rather than by the next invoice.
How a credential is verified, read from our own auth path
Credential handling is where a security page stops being a policy discussion. Ours is two code paths,
read on 2026-10-07 from backend/smartgate/core/auth.py.
- The key is never stored, only its digest. An incoming bearer token is hashed and looked up by that hash; the row is filtered to keys that have not been revoked. The practical consequence is that revoking a credential is a single flag in the database with no cache for an attacker to wait out — and that a leaked database does not hand over usable keys.
- Unknown and revoked keys produce the same refusal, with the same wording. That is deliberate: an error that distinguishes them is an oracle telling a caller which keys once existed. An API that says "no such key" for one case and "revoked" for the other has published its key inventory.
- There is a second, internal path, and it is worth naming. A trusted header carrying a team identifier is accepted from the platform's own server-side bridge, and it resolves to a team with no key id attached. That is a reasonable design for a first-party caller and a poor one for anything reachable from the internet: the security boundary becomes network position rather than a secret. If you are reviewing this pattern, the question to ask is who can reach that path, not whether the code is correct.
- Revocation, rotation and audit attribution are separate concerns. Because the record keeps the key id alongside the team, you can answer "which credential acted" after a rotation — the difference between a security review with an answer and one with a shrug.
Where SmartGate fits
SmartGate is the control plane described above, made concrete: an MCP-native algorithm gateway for token control, traffic shaping and agent audit. It is a single authenticated endpoint through which an agent reaches its tools, with per-key metering and an audit record written as calls happen. Seven tools are exposed through it — smart_fetch, smart_search, smart_context_gate, smart_dedup, smart_budget_guard, smart_memory and smart_pipe — and every call is counted against the caller's key.
It is worth being precise about what this gives a platform team and what it does not. SmartGate does not claim to detect or block prompt injection or jailbreaking, and nothing on this page should be read as saying that it does. Detection of those attacks is a model-side and application-side problem, and the pages that cover it are the attack pages in this cluster. What the gateway contributes is the columns a platform engineer can actually operate: limits that bound what a looping or misbehaving agent can spend, traffic shaping that keeps one noisy caller from affecting the others, and a record that names the key behind each call — the attribution row at the bottom of the checklist and the one an incident review asks for first.
The surfaces are documented rather than described here. Call limits live on token control, the handling of bursts is the traffic customs surface, and the record is audit and compliance. The plan table determines the operational numbers rather than the feature set, so what matters for a compliance decision is the current table on the pricing page, read against your own retention and attribution requirements.
How to get started
Each step below makes the next one cheaper, and none of them requires the model to change.
- Write the key register. List every key, its owner, and what it is allowed to reach; anything you cannot name is a control you do not have.
- Put a limit on before the volume arrives. A rate limit and a per-team quota cost nothing to set and bound the worst case from the first day; a limit added after an incident is a limit that already failed.
- Decide the log policy in writing. Metadata by default, payloads on exception, a retention window per class, and a redaction step you have tested. Publish it, because the policy is the control.
- Draw the three boundaries. Input, provider and output, each in one sentence, with the data classes each one is allowed to carry.
- Add control tests to the infrastructure suite. Quota, rate limit, revocation, redaction and attribution, each a test that fails loudly if a policy drifts.
- Route the calls through a gateway and watch one end to end. Start free and follow a single call — key, limit, record — before trusting the numbers for real volume.
Frequently Asked Questions
Does a gateway prevent prompt injection?
No. A gateway enforces deterministic limits and writes a record; it does not tell a legitimate instruction from an injected one, and claiming otherwise would overstate what any control plane can do. What a gateway gives you is a bound on the consequence, through rate limits, quotas and attribution, not a detector for the attack itself.
Should we log full prompts and responses?
Default to no. Log the metadata of every call, meaning caller, tool, tokens, latency and outcome, and keep payload bodies only when a specific purpose requires them, with a retention window, redaction and encryption attached. If a prompt can contain personal data, storing it without a policy is a disclosure waiting to happen.
One key for all our services, or many?
Many, at least one per service and environment. A shared key is a shared identity: you cannot attribute spend, revoke a single caller, or set a limit that holds for one team without affecting all of them. Per-caller keys are the foundation every other control reads from.
How often should keys be rotated?
On a schedule you can actually keep, and immediately on suspicion. The schedule matters less than the revocation path being fast and rehearsed, because the useful metric is how long a revoke takes under pressure, not how many days the key lived.
Where should a platform team start?
With the register and the limits: name every key and its owner, then put a rate limit and a quota on each before real volume arrives. Both are cheap, both are deterministic, and both make every later control easier to reason about.
Is model-level hardening enough on its own?
No. Model hardening lowers the probability of a bad outcome and does not reach zero, and it knows nothing about which data your deployment reaches or which tools it holds. It is one layer; the deterministic controls are the ones that bound what a compromise actually costs.
Limitations
This page is an operations framework, and it is deliberately narrow: it does not explain how prompt injection, indirect injection or jailbreaking work, and it makes no claim that any control here prevents them. Those mechanisms belong to the attack-focused pages in the cluster, and a platform that treated a quota as a barrier against an attacker would be misreading this page.
The controls described bound consequence and improve attribution; they do not make an untrusted system safe, and a determined attacker who can already act inside your boundary is not stopped by a rate limit. The checklist is a starting inventory rather than a compliance certification, and where a specific regulation or contract sets a retention or residency requirement, that requirement — not this page — decides the policy. Product figures quoted above are plan limits rather than feature claims and change with the tier, so the pricing page is the authoritative table.
Sources
-
OWASP — the Top 10 for LLM Applications and the Secrets Management Cheat Sheet.
-
OWASP — the Logging Cheat Sheet and the API Security Top 10.
-
NIST — the AI Risk Management Framework and its Generative AI Profile (NIST AI 600-1).
-
NIST — SP 800-207, Zero Trust Architecture, for the least-privilege and per-request verification model the gateway controls implement.
-
Model Context Protocol — the specification's security best practices.
-
Product behaviour and the plan table were read read-only from the product source at the revision this project's pipeline run recorded, with the plan figures re-checked against the live pricing page on 2026-10-04. The demand figures are this project's own paid measurement, recorded in its
search_volume.jsonandresearch_brief.md. -
The section "How a credential is verified, read from our own auth path" is our own implementation, read on 2026-10-07 from
backend/smartgate/core/auth.py(origin/main). It states only what those files state.
Method note
This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned 0 of 7 sections for this page (0 abstention(s), 7 no-slice verdict(s)): rule A found no unique symbol in the scanned repository for any of the seven section keywords, because this lane's vocabulary — keys, logs, limits and audit — collides with generic helper and configuration names across a codebase. A pinned generic name would have given the page the shape of a verified article with none of the substance, so every section but one above is written from external, linkable sources, which is the house rule for an unpinned section.
Product facts were read read-only from the product source at the revision this project's pipeline run recorded, and the plan figures were re-checked against the live pricing page on 2026-10-04. The external statements above are each source's own published wording, read at the linked pages, and no code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.