SmartGateSmartGate

Prompt Security for AI Applications: Policy, Keys, Audit

secure prompt handling is not one filter in front of a model. It is five decisions that each have to survive a bad request: parse the policy on every read, clamp it to what the plan allows, validate the API key before the prompt is accepted, mask the payload before it reaches a log, and emit one audit record a human can read.

Short answer: secure prompt handling is not one filter in front of a model. It is five decisions that each have to survive a bad request: parse the policy on every read, clamp it to what the plan allows, validate the API key before the prompt is accepted, mask the payload before it reaches a log, and emit one audit record a human can read. The code below is the shipped version of those decisions in the SmartGate gateway. prompt security carries 1,000 US searches a month; this is the verifiable part of that demand.

Key takeaways

  • Five layers, one function each. Policy parse, plan clamp, edit assertion, partial merge, key validation, masking, param parsing, audit summary - named functions you can review in one sitting, not a framework.
  • The policy parser never returns raw input. Null, undefined and the empty string normalise to an empty object, and a failed validation falls back to the schema's own defaults.
  • The plan clamp can only tighten. Minimum for rate, the plan's floor for the daily cap, and two governance switches that can only be turned off.
  • The key check fails closed on three conditions: no row for the key hash, a revoked row, an expiry in the past.
  • Masking is name-based and recursive: api_key, secret and token become three asterisks at any depth.
  • The audit line is legible, not safe. A fixed field order with the URL truncated at 48 characters, and the URL and query written verbatim.
  • Do this next: implement the key check this week - hash the key, look it up by hash, refuse revoked and expired rows - before you write a policy you cannot enforce.

The short version for whoever signs off on the risk

"Prompt security" sounds like a filtering problem, and filtering is the expensive half: a model that reads every prompt before the model itself does, tuned forever. The cheaper half is governance, and it is the half most teams skip. Governance means four questions have an answer in production: what is this team allowed to do, who is calling right now, what may end up in a log, and what record do we keep.

The code here is the shipped answer to those questions in the SmartGate gateway - an MCP-native algorithm gateway for token control, traffic shaping, and agent audit. One policy object per team, validated on every read; a clamp that only reduces what the plan allows; a key check that refuses revoked and expired credentials; masking applied recursively by key name; an audit line assembled from a fixed field order. None of it is a classifier, and all of it is reviewable in an afternoon.

Pricing follows the platform rather than the number of decisions: Free is $0 with 2M tokens a month, Pro starts at $18, Teams at $55, Enterprise is contract pricing, and billing is "pay for the platform, share only when you save" - capped around $36 a month on Pro. The platform slogan is the reason the layer exists at all: agent loops don't warn, they bill.

What the prompt-security market actually measures

Two measured phrases frame this page. prompt security takes 1,000 searches a month in the US with a flat 12-month range of 880 to 1,600. prompt injection protection takes 50 a month with a keyword difficulty of 39, trending up from a floor of 30. Both are Google Ads figures for US English, pulled on 2026-09-15.

Their intent differs, and the honest read matters. prompt security is dominated by a company that happens to be named Prompt Security: the top organic result is its own site, the third is its Wikipedia entry, a knowledge graph sits on the page, and all four People Also Ask questions are about the company. There is no AI Overview. A page on this phrase competes with a brand rather than a technique, and should expect a fraction of those 1,000 searches.

prompt injection protection is the opposite: it carries an AI Overview, and that overview is a list of named layers - contextual separation, input validation, output filtering, least privilege, human in the loop - citing the OWASP prompt-injection cheat sheet. The category's frame is already established by OWASP's LLM01 entry in the Top 10 for LLM Applications, the NIST AI Risk Management Framework, Microsoft's AI red team guidance, Google's Model Armor, and Anthropic's jailbreak mitigation guidance.

Those named layers describe the runtime surface around a call — tool permissions, MCP server trust, and how untrusted content is marked — worked through in depth on AI agent security.

What none of them give you is the per-request plumbing under the layer names: which policy was in force, whether the caller's key was still valid, what the log recorded. That is what the functions below do.


parseGatewayPolicy: normalise first, then trust the schema

# lib/settings/schemas/gateway-policy.v1.ts — source lines 35–42 (parseGatewayPolicy)
function parseGatewayPolicy(raw: unknown): GatewayPolicyV1 {
  const empty = raw === null || raw === undefined || raw === "" ? {} : raw;
  const parsed = gatewayPolicyV1Schema.safeParse(empty);
  if (parsed.success) {
    return parsed.data;
  }
  return gatewayPolicyV1Schema.parse({});
}

The input is typed unknown because it comes from a stored column no type system protects. Null, undefined and the empty string all collapse to an empty object before validation. The schema then validates safely: on success the parsed data is returned, so an omitted field arrives as the schema's default. The failure branch is the interesting one: it runs the schema's throwing parse on an empty object instead of returning the raw value. That either produces the platform defaults or raises, and a corrupt policy never becomes an absent one.

That versioned schema is also the contract a consumer reads, and keeping the published reference true to it is the practice set on API documentation best practices for AI tools.

loadGatewayPolicy: one read path, three lines

# lib/settings/load-gateway-policy.ts — source lines 10–12 (loadGatewayPolicy)
function loadGatewayPolicy(team: TeamPolicySource): GatewayPolicyV1 {
  return parseGatewayPolicy(team.gatewayPolicy);
}

Three lines buy one property: there is exactly one place where a team's stored policy becomes an object, and everything else goes through it. The alternative - each service reading the column and interpreting it locally - is how two components end up enforcing two policies for the same team.

clampGatewayPolicyToPlan: the plan can only tighten the policy

# lib/settings/clamp-gateway-policy.ts — source lines 4–35 (clampGatewayPolicyToPlan)
function clampGatewayPolicyToPlan(
  policy: GatewayPolicyV1,
  caps: PlanCapabilities,
): GatewayPolicyV1 {
  const rpm = Math.min(policy.playground.rpm, caps.maxPlaygroundRpm);

  let dailyCap = policy.playground.daily_request_cap;
  if (caps.minPlaygroundDailyCap != null) {
    dailyCap = caps.minPlaygroundDailyCap;
  } else if (
    dailyCap != null &&
    caps.maxPlaygroundDailyCap != null
  ) {
    dailyCap = Math.min(dailyCap, caps.maxPlaygroundDailyCap);
  }

  return {
    ...policy,
    playground: {
      rpm,
      daily_request_cap: dailyCap,
    },
    governance: {
      allow_member_tool_overrides: caps.memberToolOverridesAllowed
        ? policy.governance.allow_member_tool_overrides
        : false,
      allow_integrator_hmac_budget: caps.l2HmacBudgetAllowed
        ? policy.governance.allow_integrator_hmac_budget
        : false,
    },
  };
}

The plan's capability object is a ceiling and this function moves the policy down to it. Requests per minute becomes the minimum of policy and plan. The daily cap is a two-branch decision worth reading twice: when the plan declares a minimum daily cap that value is assigned outright, so a plan floor overrides a higher policy value; only when no floor exists and both sides declare a maximum does the lower win. The governance switches are strictest of all - member tool overrides and integrator HMAC budgets can only be turned off here, never on.

Where those tightened limits are counted, and which layer enforces the per-key rate and the per-team budget instead of every service re-deriving them, is the control plane on AI agent governance.

assertGatewayPolicyEditable: the write path refuses before it writes

# lib/settings/assert-plan-allows.ts — source lines 40–44 (assertGatewayPolicyEditable)
function assertGatewayPolicyEditable(caps: PlanCapabilities) {
  if (!caps.teamGatewayPolicyEditable) {
    throw new PlanFeatureDisabledError("PRO", "gateway_policy");
  }
}

Where the clamp quietly shapes a value, this function throws: if the plan capability says a team's gateway policy is not editable, the caller gets an error carrying both the plan that unlocks the feature and the feature name. Two choices are worth copying. It is an assertion rather than a boolean, so a caller cannot forget to check it. And it names the feature, so an API can answer "not on your plan" with the plan that would change that, instead of a bare forbidden.

mergeGatewayPolicyPartial: five sections, shallow on purpose

# lib/settings/schemas/gateway-policy.v1.ts — source lines 44–55 (mergeGatewayPolicyPartial)
function mergeGatewayPolicyPartial(
  current: GatewayPolicyV1,
  partial: Partial<GatewayPolicyV1>,
): GatewayPolicyV1 {
  return gatewayPolicyV1Schema.parse({
    defaults: { ...current.defaults, ...partial.defaults },
    playground: { ...current.playground, ...partial.playground },
    governance: { ...current.governance, ...partial.governance },
    analytics: { ...current.analytics, ...partial.analytics },
    aut_baseline: { ...current.aut_baseline, ...partial.aut_baseline },
  });
}

Partial updates are where configuration rots: a caller sends the one field it wanted to change and quietly drops the rest. This function merges five named sections - defaults, playground, governance, analytics and the autonomy baseline - then re-validates the result through the same schema that guards a full write. Two honest qualifications. The merge is shallow per section, so a nested object is replaced wholesale and a caller wanting one field inside it must send the whole section.

isEmptyPolicy: an empty policy is a decision, not a blank

# lib/settings/plan-seeds.ts — source lines 15–21 (isEmptyPolicy)
function isEmptyPolicy(raw: unknown): boolean {
  if (raw === null || raw === undefined) return true;
  if (typeof raw === "object" && Object.keys(raw as object).length === 0) {
    return true;
  }
  return false;
}

This test decides when a stored plan seed should have defaults written to it, and it is deliberately narrow rather than a validator: null and undefined are empty, an object with no own keys is empty, everything else is not. One edge case: because the type check is on "object" rather than on a plain object, an empty array also reads as empty. A junk string value is not empty, so it goes to the policy parser and fails there, loudly, which is the safer failure for this caller.


validateApiKey: three reasons to refuse, one write per call

# lib/api-keys/index.ts — source lines 92–118 (validateApiKey)
async function validateApiKey(
  rawKey: string
): Promise<{ valid: boolean; teamId?: string; keyId?: string }> {
  const hash = hashKey(rawKey);
  const key = await prisma.apiKey.findUnique({
    where: { keyHash: hash },
    select: {
      id: true,
      teamId: true,
      revokedAt: true,
      expiresAt: true,
    },
  });

  if (!key || key.revokedAt) return { valid: false };
  if (key.expiresAt && key.expiresAt < new Date()) return { valid: false };

  await prisma.apiKey.update({
    where: { id: key.id },
    data: {
      lastUsedAt: new Date(),
      usageCount: { increment: 1 },
    },
  });

  return { valid: true, teamId: key.teamId, keyId: key.id };
}

Every prompt through a gateway arrives with a key, and this is where it is judged. The function hashes the raw key and looks up by hash, so the plaintext credential is never the lookup column, and it selects four fields: row id, team id, revocation timestamp, expiry timestamp. Three conditions return an invalid result with nothing attached - no row, a revoked row, an expiry in the past. Only after those pass does it write, marking the key used and incrementing a usage counter, so an invalid key costs one read and no write. Validity is a boolean, so a caller explaining a refusal must derive the reason itself - and the update sits on the request path, which is why a request-scoped cache belongs above it.

cfCredentials: credentials from the environment, or an empty string

# lib/pseo/kv-read.ts — source lines 3–9 (cfCredentials)
function cfCredentials() {
  return {
    accountId: process.env.CLOUDFLARE_ACCOUNT_ID ?? "",
    namespaceId: process.env.CLOUDFLARE_KV_NAMESPACE_ID ?? "",
    apiToken: process.env.CLOUDFLARE_API_TOKEN ?? "",
  };
}

Three environment variables, three empty-string defaults, one object. There is no hardcoded fallback token and no default namespace identifier, so a deployment missing a variable cannot silently address a different account - it fails on the first call. The file declares itself server-only, so a bundler refuses to pull the credentials into client code.

mask_sensitive: redact on the way into the log

# backend/smartgate/shared/utils.py — source lines 9–21 (mask_sensitive)
def mask_sensitive(data: Dict[str, Any], keys: set = {"api_key", "secret", "token"}) -> Dict[str, Any]:
    """脱敏敏感字段."""
    result = {}
    for k, v in data.items():
        if k in keys:
            result[k] = "***"
        elif isinstance(v, dict):
            result[k] = mask_sensitive(v, keys)
        elif isinstance(v, list):
            result[k] = [mask_sensitive(i, keys) if isinstance(i, dict) else i for i in v]
        else:
            result[k] = v
    return result

The rule set is a default parameter: three key names - api_key, secret and token - become three asterisks, and the function recurses into nested objects and into lists whose items are objects. Lists of plain values pass through unchanged. The match is on the key name, not the value, so a secret under an unlisted name is not masked - the rule set is a contract with the writer, and it has to change when the writer does. And masking applies at any depth, which is what makes it usable on request payloads.

mask_key: eight characters of context

# backend/tools/cursor_mcp_deep_test.py — source lines 47–50 (mask_key)
def mask_key(key: str) -> str:
    if len(key) <= 12:
        return "***"
    return f"{key[:8]}...{key[-4:]}"

A different trade-off, for one value on one line. Keys of twelve characters or fewer are replaced entirely with three asterisks; a longer key keeps its first eight characters and its last four, separated by an ellipsis. That is enough to recognise which key a log line refers to and not enough to use it. Be honest about where the helper lives: a command-line deep-test harness for an MCP client, not the gateway request path. That is the point - redaction is a rule you apply to every tool that prints a credential.

parseParams: audit params that survive a bad writer

# lib/smartgate/audit-logs.ts — source lines 48–63 (parseParams)
function parseParams(raw: unknown): Record<string, unknown> {
  if (raw && typeof raw === "object" && !Array.isArray(raw)) {
    return raw as Record<string, unknown>;
  }
  if (typeof raw === "string") {
    try {
      const parsed = JSON.parse(raw) as unknown;
      if (parsed && typeof parsed === "object" && !Array.isArray(parsed)) {
        return parsed as Record<string, unknown>;
      }
    } catch {
      /* ignore */
    }
  }
  return {};
}

Audit parameters arrive as whatever the writer stored: an object, a JSON string, or garbage. A non-array object is returned as-is; a string is parsed inside a try/catch and accepted only if it is a non-array object; everything else becomes an empty object. That last line is the design decision - an unreadable payload produces a row saying the parameters could not be read instead of dropping the event or throwing inside logging. Arrays are rejected on purpose, because treating a list as a record would let a positional index become a field name.

buildSummary: one line a human can act on

# lib/smartgate/audit-logs.ts — source lines 72–93 (buildSummary)
function buildSummary(
  auditTool: string,
  params: Record<string, unknown>,
  tokenUsed: number | null,
  success: boolean,
): string {
  const parts: string[] = [];
  const route = params.route as string | undefined;
  if (route) parts.push(`route=${route}`);
  if (params.url && typeof params.url === "string") {
    parts.push(params.url.length > 48 ? `${params.url.slice(0, 48)}…` : params.url);
  }
  if (params.query && typeof params.query === "string") {
    parts.push(`q=${params.query}`);
  }
  if (tokenUsed != null && tokenUsed > 0) {
    parts.push(`${tokenUsed.toLocaleString()} tokens`);
  }
  if (!success) parts.push("failed");
  if (parts.length === 0) return auditTool || "audit";
  return parts.join(" · ");
}

The summary is assembled in a fixed order - route, URL, query, token count, failure marker - and the order is the contract, because the first thing in the line is where the request went. The URL is truncated at 48 characters with an ellipsis so one long parameter cannot push everything else out of view; the token count is locale-formatted and appears only above zero; failures append the word failed; an empty part list falls back to the tool name. The honest wrinkle: the URL and the query string are the two fields a prompt can influence, and both are written verbatim. That is why masking belongs upstream - the audit summary is built to be legible, not safe.

A row this legible is also the input a report is built from, which is the path AI Financial Analysis for Agents takes from audit log to analysis.


How SmartGate compares

The comparison is about where policy, identity and evidence live - not about which vendor has the better detector.

What it actually enforces Where the policy lives What you pay
Do nothing, or hand-roll it in the app Whatever each service remembers to check; records nothing independently Scattered across services and repositories Engineer time, plus the incident nobody can reconstruct
Model-provider guardrails Some classes of unsafe content, at the provider boundary In the provider's console, per provider In the model bill, duplicated per provider
A dedicated prompt-security scanner Detection of adversarial input, with a model of its own In the vendor's console A separate subscription, and a second place your prompts travel
An API router or proxy Routing and per-token accounting; policy is a config file In the proxy's configuration A per-token fee; meters, keys and masking are still yours
Gateway with policy, key and audit primitives (SmartGate) A schema-validated policy per team, a clamp that only tightens, per-request key validation, masked audit rows In one object the gateway reads on every request Platform fee, plus a share only once measured savings pass a threshold

The sentence to test against your own bill is pay for the platform, share only when you save: on Pro the share is capped around $36 a month, nothing is charged while measured savings are below the threshold. Free is $0 with 2M tokens a month and 7 days of logs; Pro starts at $18 a month with 20M tokens and 30 days; Teams at $55 with 100M tokens and 90 days; Enterprise is contract pricing, all on the pricing page.

How to get started

  1. Put the policy in one object. One schema, one version, one validating read function that every service calls.
  2. Clamp it on write. Minimum for rate, floor for the daily cap, governance switches that can only be turned off, so stored policy is never illegal for the plan.
  3. Check the key before the prompt. Hash, look up by hash, refuse revoked and expired rows early, and count the use only after the checks pass.
  4. Mask before you log. Start with three key names applied recursively, and treat the rule set as a contract that changes whenever the writer changes.
  5. Write one line per call. Fixed field order, truncated URL, token count, failure marker.
  6. Leave detection to the detector. Layer this underneath a scanner or a guardrail model, not instead of one.

Start on the free tier - 2M tokens a month, all seven smart_* tools (smart_fetch, smart_search, smart_context_gate, smart_dedup, smart_budget_guard, smart_memory, smart_pipe), no card: start free. The tool surface is in the docs, the audit path in audit and compliance, and the plan limits on the pricing page.

The tools those keys unlock are what an individual wires into their own assistant, described from the user's side on SmartGate MCP for research and decisions.

Frequently Asked Questions

Is masking by key name good enough?

Name-based masking catches everything stored under a name you listed, at any depth, which covers what your own writers produce. It does not catch a secret pasted into free text or stored under an unlisted name, so pair it with a value-pattern check.

Should the plan clamp run on read or on write?

On write. Clamping on write means stored policy is always legal for the plan and every reader can trust it without re-deriving the limits. Clamping on read is the fallback for values that arrived from a migration or a direct database edit - a safety net, not the primary control.

Does validating a key on every request cost too much?

One indexed lookup by hash plus one counter update per request is cheap next to a model call. The thing to watch is freshness: the write only happens after the three refusal conditions pass.

Can this stop prompt injection?

No, and it is important to say so. These functions decide what a team may do, who is calling, and what the record shows. Injection defence needs contextual separation, input validation and output filtering around the model. Governance makes those layers accountable; it does not replace them.

What happens if the stored policy is corrupted?

The parser normalises the three empty shapes and then validates. If validation fails, the failure path resolves to the schema's defaults or raises - it never returns the unvalidated value, so a corrupt policy degrades to a known-good default rather than to whatever was in the column.

Do we still need a scanner in front?

If you have one, keep it. A scanner inspects content; this layer records identity, policy and outcome. They answer different questions after an incident, and the governance record is the one that says which team, which key and which policy were in force at the time.

Limitations and what this does not do

  • None of these layers is a detector. There is no classifier here and no prompt-content inspection. The code decides policy, identity, log content and audit shape; it does not decide whether a prompt is malicious.
  • Masking is name-based, so the rule set is the whole guarantee. A secret under an unlisted key name, or pasted into free text, reaches the log unchanged.
  • The audit summary writes the URL and the query verbatim. That is a legibility choice with a privacy cost, and it is why the masking step has to sit upstream of it.
  • Retention is a plan property, not a design property. 7 days on Free, 30 on Pro, 90 on Teams, 180 on Enterprise. If your regulatory position needs longer, your own archive is what you must add. How a retention period is chosen from the duty that applies to you, and how the record is exported to someone who did not build it, is the evidence discipline on AI compliance.
  • Failure behaviour is not uniform across layers. The parser defaults, the clamp tightens silently, the edit assertion throws. A control that must fail loudly needs the assertion treatment.
  • Two helpers sit outside the request path. The key-masking and credential helpers come from a test harness and a content-read path, not from the prompt pipeline.

This page owns the request path, not the programme around it: the policy layers, their owners and their review cadence are the subject of the AI governance framework, and a control that is not placed in one of those layers will be re-litigated every quarter.

Method note

The code in this article is not transcribed. Each block was cut directly out of the slice body returned by the SmartGate slice API and re-asserted byte-for-byte as a substring of that body; the first line inside every fence records the file and the exact source lines. Symbols were pinned by whole-name containment (rule A level 2) and confirmed by the service's slot-proof endpoint before being written into the prose - 12 of 12 planned sections pinned, no abstentions, no misses. All twelve blocks are complete function bodies: no window into a slice was needed, because the longest is 32 lines.

The demand figures come from a separate DataForSEO Google Ads volume call recorded in work/batch6/, not from this project's own plan run: all twelve of this project's seed phrases returned no advertiser data, which is a seed-design signal rather than a connectivity problem. The prompt security and prompt injection protection numbers are therefore cited as borrowed cluster demand rather than this page's own.

Slice provenance

# SERP keyword Symbol File Source lines How it was pinned sha256(12)
1 parseGatewayPolicy parse gateway policy for prompt security parseGatewayPolicy lib/settings/schemas/gateway-policy.v1.ts 35–42 rule A L2 → slot-proof dd074c32468e
2 loadGatewayPolicy load the gateway policy for secure prompt handling loadGatewayPolicy lib/settings/load-gateway-policy.ts 10–12 rule A L2 → slot-proof 82918a7047cc
3 clampGatewayPolicyToPlan clamp gateway policy to the plan limits clampGatewayPolicyToPlan lib/settings/clamp-gateway-policy.ts 4–35 rule A L2 → slot-proof 613facc8545b
4 assertGatewayPolicyEditable assert gateway policy editable before changes assertGatewayPolicyEditable lib/settings/assert-plan-allows.ts 40–44 rule A L2 → slot-proof 5bf979730866
5 mergeGatewayPolicyPartial merge gateway policy partial updates safely mergeGatewayPolicyPartial lib/settings/schemas/gateway-policy.v1.ts 44–55 rule A L2 → slot-proof b35e63b73ffa
6 isEmptyPolicy is empty policy handling for secure defaults isEmptyPolicy lib/settings/plan-seeds.ts 15–21 rule A L2 → slot-proof 142457b5325c
7 validateApiKey validate api key before accepting prompts validateApiKey lib/api-keys/index.ts 92–118 rule A L2 → slot-proof eea805141bb6
8 cfCredentials cf credentials handling for prompt security cfCredentials lib/pseo/kv-read.ts 3–9 rule A L2 → slot-proof 9e851e1e5cde
9 mask_sensitive mask sensitive fields in prompt logs mask_sensitive backend/smartgate/shared/utils.py 9–21 rule A L2 → slot-proof 1776f5449b8d
10 mask_key mask api keys in prompt audit output mask_key backend/tools/cursor_mcp_deep_test.py 47–50 rule A L2 → slot-proof 1ac4b9165914
11 parseParams parse params for prompt audit logs parseParams lib/smartgate/audit-logs.ts 48–63 rule A L2 → slot-proof 747608709d38
12 buildSummary build summary of handled prompts buildSummary lib/smartgate/audit-logs.ts 72–93 rule A L2 → slot-proof 009eace0bd53

Every fenced block above was cut from the slice body and re-asserted against it byte-for-byte before publication. 12 of 12 sections pinned, 0 abstentions, 0 misses.