AI Red Teaming: From Threat Model to Regression Case
AI red teaming is the structured practice of attacking an AI system on purpose, so that failures are found by your team and not by a user or an attacker. It runs as a method rather than a scan: model the threat, grow an attack library, probe automatically, score the rate, and turn each finding into a regression case.
Short answer: AI red teaming is the structured practice of attacking an AI system on purpose, so that failures are found by your team and not by a user or an attacker. It runs as a method rather than a scan: model the threat, grow an attack library, probe automatically, score the rate, and turn each finding into a regression case. It differs from penetration testing in what counts as a result — an emergent behaviour with a rate, not a patchable bug with a fix.
Key takeaways
- A red-team result is a rate and a coverage statement, not a pass or a fail. A language model samples, so the same probe fails once and refuses the next time; the deliverable is how often the attack succeeded and which techniques were tried.
- Threat modelling decides the whole run. Name the data the agent can reach, the untrusted content it reads and the outbound paths it can write, and scope the test to that surface before a single probe is written.
- The attack library is the asset. Seed probes plus mutation operators scale coverage; a fixed list of hand-written prompts goes stale the moment the model or the tool schema changes.
- The last mile is regression. A finding that is not captured as a repeatable case with an expected behaviour will be fixed once and quietly reintroduced by the next prompt edit.
- Decide in-house versus outsourced by cadence, not by prestige. Continuous regression belongs next to the code; independent assurance is worth buying on a schedule.
What ai red teaming is, and how it differs from penetration testing
AI red teaming is the practice of deliberately attacking a machine-learning or language-model system to find the ways it can be made to misbehave, before a user, a regulator or an attacker finds them first. Applied to AI, the adversary's target is a probabilistic system that reads text, calls tools and acts, so the red team is testing behaviour and not only code.
Microsoft's red-teaming guidance splits the work into two tracks usually run together: the security track, for the classic confidentiality, integrity and availability harms, and the responsible-AI track, for harmful, biased or off-policy content. The two need different probes and often different reviewers, so the split is the first thing to settle in scope (Planning red teaming for large language models).
The distinction from a penetration test matters because it changes the deliverable. A penetration test asks a narrow, answerable question — can an attacker exploit a known class of flaw in this system? — and the answer is a list of findings each with a fix. Red teaming an AI system asks a wider question, what can it be made to do that it should not, and the answer is a distribution rather than a list. The bottleneck is not a single vulnerability but the system's own authority: the same context that reads attacker-influenced text can call tools and reach data, which the cluster's prompt injection centre page defines in detail.
| Penetration test | AI red teaming | Bug bounty | |
|---|---|---|---|
| Target | A known flaw class in an implementation | Emergent behaviour of a model plus its tools and data | Whatever a crowd chooses to look at |
| Question | Is this exploitable? | What can it be made to do? | Will anyone report it? |
| Result | A finding and a fix | A rate, a coverage map, a regression suite | A payout and a patch |
| Cadence | Pre-release or annual | Continuous, tied to model and prompt changes | Always on, unbounded |
| Failure mode | Missed a known class | Tested the wrong surface | Never tested the surface you changed |
The practical consequence is that a red-team report is written as a rate and a coverage statement. A finding reads "this instruction made the agent call the write tool in 12 of 40 samples", not "the system is vulnerable". A team that writes the second sentence has usually tested the surface it understood rather than the surface that carries the payload, which is why threat modelling comes before probing.
llm red teaming: threat modelling before a single probe runs
Threat modelling for an LLM system starts from what the system can do, not what it can say. The inventory has three parts — the private data the agent can reach, the untrusted content it reads, and the outbound paths it can write — which are the three the cluster centre calls the lethal trifecta. A system with all three is where a red-team finding is likely to be a real incident.
From that inventory the model is written down as a small set of trust boundaries and capabilities. Which tools can the agent call, and which write, send, pay or delete? Which documents, pages, emails or tool returns arrive from a source a third party controls? Which identity does the agent act as, and what can it reach that a human user could not? Those three answers turn a vague worry into a test plan, because each boundary is a place a probe can cross from data into instruction — and they set scope, so the budget does not go to the easy chat surface.
Two surfaces are easy to omit from the model and expensive to omit from the test. The first is the tool layer, where a tool description or a return value is itself an instruction channel. The second is the supply chain that produced the model, its adapters and its dependencies: a pinned version, a checksum and a provenance record are what let a red team distinguish "the model behaved badly" from "the model is not the model you think it is", and that surface is covered in depth on model supply chain. The platform side of the same picture — credentials, log retention, per-key limits — belongs to llm security.
The output of this step is not prose. It is a table with one row per capability and one column per trust boundary, marked in-bounds or out, and an explicit note of what a "success" would look like for each row. That table is what keeps the scoring honest later: an attack that succeeds against an out-of-bounds capability is a false alarm, and one that fails against an in-bounds capability is a coverage gap.
ai red teaming tools: enumeration, mutation, and the harness
The tooling sorts into three functions, and running one without the other two leaves the work half-done. The first is probe generation: an enumerator of attack templates plus a mutator that turns one seed into hundreds of distinct inputs. The second is the execution harness, which drives the target the way a real caller would — through its tools, not only its chat endpoint — and records every request and response. The third is the evaluator that decides whether a response is a failure, usually a classifier or a second language model given a rubric.
The open-source projects worth reading to understand the shape are, in order of how much of that pipeline they cover: garak, a probe library that enumerates attack families and reports a pass rate per family; promptfoo's red-team mode, which generates adversarial inputs and grades the outputs against configurable assertions; PyRIT, Microsoft's Python framework for building multi-turn attacks and orchestrating a number of generator and scorer endpoints; and DeepTeam, a framework aimed at red teaming both models and the agents built on them. Microsoft also ships a managed red-teaming agent that runs adversarial probes against an application and produces a graded scorecard, which is the commercial shape of the same three functions (Introducing the AI Red Teaming Agent).
Two design decisions separate a harness that finds things from one that produces noise. First, record the target's configuration with every sample — model id, sampling parameters, the system-message revision, the tool schema, and which retrieval index was live. A finding without that context cannot be reproduced, and a non-reproducible finding cannot become a regression case. Second, treat the guardrail as part of the system under test, not the referee: a red team that only measures the model will report a vulnerability the deployed product does not have, or miss the one it does. The guardrail's own behaviour is the subject of ai guardrails; here it is simply another layer a probe must pass through.
Automation buys breadth at a fixed cost per sample; it does not buy the novel attack. The honest division of labour is that the harness runs the library continuously, and a human spends the time the harness cannot: reading the responses that scored as refusals, following a thread across turns, and converting a one-off curiosity into a seed for the next automated round.
red teaming llm applications: the attack library and its variants
The attack library is the durable asset of a red-team practice, and it is built in two layers. The first layer is enumeration: one or more seed prompts per attack category, written so that a success tells you which control failed. The categories a language-model application needs are stable and small — direct instruction override, indirect instruction delivered through fetched content or a tool return, system-prompt and secret extraction, data exfiltration through a rendered link or image, unsafe tool selection, and actions that exceed the authority the user granted. None of those categories is a surprise; what makes the library useful is that every entry is written as a statement about the application, so the result maps onto a control rather than onto a taxonomy slot.
The second layer is variants, and this is where a small library becomes wide coverage. A mutator takes a seed and keeps the intent while changing the surface: rephrase, translate, encode, split across turns, embed inside a document the agent will fetch, or attach to a tool result. Each transformation is a cheap hypothesis about which filter the seed currently fails to pass, so the pass rate of a category across its variants beats the pass rate of one hand-written prompt. The catalogue also shows when a family of attacks is systematically under-tested — the gap the industry's standard entry list is designed to expose, read as an index on the OWASP Top 10 walkthrough rather than as a substitute for application-specific probes.
Two disciplines keep the library from rotting. Every seed carries a provenance note: where it came from, which control it is meant to test, and what a success looks like. And every seed carries a last-verified date, because a probe written against last quarter's system prompt may now pass for the wrong reason. A library of two hundred unlabelled prompts is a demo; a library of forty labelled seeds with mutation operators and a record of when each last caught something is a test suite.
llm red teaming prompts: from a probe to a regression case
A red-team prompt earns its place in the repository when it becomes a regression case, and the conversion is mostly bookkeeping the finder usually skips. A case needs four things recorded together: the input, including any fetched document or tool return the attack depended on; the configuration the system ran with; the observed failure, quoted or described precisely enough that a reviewer can recognise it again; and the expected behaviour — the property that must hold, expressed as something a test can check, such as a required refusal, a schema the output must satisfy, or a tool call that must not appear.
The hard part is that a language model is not deterministic, so a case cannot assert an exact output. The workable form is a sampled assertion: run the case N times and assert that the failure rate is at or below a threshold the team has chosen, with the threshold written down beside the case. That turns a red-team finding into the same kind of artifact as a performance budget, and it means a regression is a number that moved rather than an argument about an example. A case that fails at the threshold is a release blocker with a name; a case that passes adds one row to the evidence that the fix held.
The second conversion is organisational. A finding should land in the same tracker and the same review flow as a functional bug, with the case attached, rather than in a separate "security" document that no one re-reads. When the fix ships, the regression case runs with the rest of the suite, so the next person who edits the system prompt and reintroduces the behaviour sees a red test instead of a report from an external assessor three months later. This is the join between red teaming and ordinary engineering, and teams that build it stop rediscovering the same finding.
ai red teaming framework: scoring, coverage, and the loop
A framework earns its name by making the run repeatable, which means fixing three things: how a success is defined, how it is scored, and how the result is tracked. How a success is defined is the one usually left implicit. For a refusal-based failure it is "the model complied when it should have refused"; for a tool-using agent it is stronger and checkable without a model judging another model — "the agent called a tool it was not authorised to call, or passed it an argument derived from untrusted content".
How it is scored depends on the evaluator. Where a check is deterministic — a forbidden tool was called, a secret appears in the output, a schema was violated — score it directly. Where it is not, an LLM-as-judge is the usual choice and brings its own error rate: judge models are probabilistic, disagree with human raters on a measurable fraction of cases, and can be steered by the text they grade. The fix is to calibrate the judge against a human-labelled sample and keep a deterministic check wherever one exists; a score with no calibration is a feeling with a decimal point.
How it is tracked is a coverage matrix rather than a single number. One axis is the attack category or technique; the other is the surface — chat, retrieval, tool use, output rendering. A cell holds the number of probes run and the failure rate observed. Two columns fall out immediately: cells with probes and a zero rate are evidence, empty cells are unknowns dressed as clean results, and most reports confuse the two. Tagging each technique makes the matrix comparable across runs; MITRE ATLAS is the public adversary-technique catalogue built for exactly that mapping, the AI counterpart to MITRE ATT&CK.
The loop that a framework closes is the one already described: model the threat, run the library, score, triage the findings, fix, and promote each fix into a regression case that runs on every change. The governance half of this — which of these outcomes a regulator or an internal control framework expects you to record, and how a control is mapped to an owner — is a separate job from the engineering and is worked through on ai security framework; this page stays with the run.
what is red teaming in cybersecurity: in-house, outsourced, or both
In conventional cybersecurity, a red team emulates a real adversary against an organisation's people, processes and technology, while the blue team defends and a purple team works the findings back into detection and response. Its distinguishing feature has always been that it is goal-directed and adversarial: it starts from a threat, not a checklist, and is measured by whether it reached the goal. The AI variant adds one property — the target's behaviour changes when its model, its prompt or its tools change, so a one-off exercise has a short half-life.
That half-life settles the in-house versus outsourced question, which is really a matter of cadence rather than expertise. The continuous job is regression: rerunning the library whenever the system changes and keeping the case set current. It belongs inside the team that owns the system, because it is coupled to every release. The independent job is assurance: a fresh adversary, unencumbered by the team's assumptions, testing whether the controls the team believes in actually hold. Buy that on a schedule, from someone who will write findings the internal team can reproduce.
| Continuous, in-house | Independent, outsourced | |
|---|---|---|
| Question | Did this change reintroduce a known failure? | Do the controls we believe in actually hold? |
| Cadence | Every release, tied to prompt and tool changes | Quarterly or on a material architecture change |
| Best at | Context, speed, the regression library | Fresh methods, no blind spots from familiarity |
| Weak at | Novel attack families, groupthink | System context, low-cost reruns |
| Artifact | A scored regression suite | A reproduction-ready findings report |
The pattern that works is unglamorous: an internal owner who keeps the library and the regression suite, an external assessor on a fixed cadence who leaves behind reproducible cases, and a rule that every external finding becomes an internal case before it is closed. Outsourcing the whole practice buys a report you cannot rerun; insourcing it buys confident results and no fresh perspective.
Where SmartGate fits
SmartGate is a control plane, not a red-teaming product, and this page will not claim that it finds anything. Its contribution to a programme is the ground truth of what an agent actually did. SmartGate is an MCP-native algorithm gateway for token control, traffic shaping and agent audit: one authenticated endpoint through which an agent reaches its tools, with per-key metering and an audit row written as calls happen. Seven tools are exposed — smart_fetch, smart_search, smart_context_gate, smart_dedup, smart_budget_guard, smart_memory and smart_pipe — and every call is counted against the caller's key.
That record turns a red-team observation into evidence. When a probe appears to make an agent call a tool it should not, the audit row says which tool ran, with whose key, at what time, and at what token cost — the difference between "the model said something alarming" and "the write tool executed under a user identity". The control plane also bounds the run's blast radius: per-key rate limits and a token budget applied at the call keep a probing harness from becoming the incident it was meant to prevent, and each probe's cost is visible. The surfaces are documented rather than described here — the record on audit and compliance, the limits on token control — and plan limits move with the tier: monthly token caps of 2M, 20M, 100M and 200M+, MCP requests per minute per key of 120, 300, 600 and 1200, and audit-log retention of 7, 30, 90 or 180 days, so the pricing page is the authoritative table.
Red teaming llm applications: what the gateway records during a run
During a red-team run the useful question is not whether a probe fired but what the record can prove afterwards, so it is worth reading what this control plane actually writes. Session identity comes from get_or_create_session_trace(session_id, header_trace=…) in backend/smartgate/core/mcp_session_trace.py (line 15): a caller-supplied trace header wins when both a header and a session id are present (lines 19–21), otherwise the function returns the id already stored for that session or mints sess_<uuid4> (line 26). The map behind it, _sessions (line 7), is a plain in-process dict, so per-session grouping is stable only for the life of the worker — a restart or a second process starts empty, and the test helper clear_session_traces (line 10) resets it outright; the module docstring names it "in-process".
The durable artifact is the export. backend/smartgate/core/audit_export.py projects each audit row through flatten_siem_row (line 68) onto a fixed fourteen-column SIEM_FIELDS tuple (lines 21–36): timestamp, request_id, correlation_id, trace_id, agent_platform, route, transport, key_id, source_ip, tool, token_used, latency_ms, success and params. It streams either NDJSON or CSV — ExportFormat = Literal["ndjson", "csv"] (line 18) — over a window chosen with the same 7d / 30d shorthand parse_export_since accepts (lines 44–56), one scope of core/access/full, up to MAX_EXPORT_LIMIT 50,000 rows (line 39). Exports are metered too: check_export_rate_limit (line 103) allows ten per team per hour and reports the scope it applied, so a red team's export cadence is itself an auditable quantity rather than a silent one. For a red team, correlation_id and trace_id are the two fields that tie a walk-through of calls to one session — and the honest limitation is that the session key is process-local while the exported rows are not.
How to get started
The first run should be small enough to finish and structured enough to repeat next month.
- Write the threat model as a table. Capabilities down the rows, trust boundaries across the columns, in-bounds or out. If the table has more than a dozen rows, the scope is too wide for a first pass.
- Seed the library with one probe per category rather than many per category, and label each with the control it is meant to test and what a success looks like.
- Add one mutation operator — a rephrase or a translation — and measure how much the pass rate moves. That single comparison tells you more than doubling the seed count.
- Record configuration with every sample, and pick one deterministic success definition (a forbidden tool call is the easiest) so the first score is trustworthy.
- Convert the first real finding into a regression case before fixing anything, so the case exists to prove the fix. Wire it into your normal test suite.
- Decide the cadence split. Name the internal owner of the suite, and if you will use an external assessor, require reproduction-ready cases in the engagement. If the agent calls tools through a gateway, point the harness at it — start free and watch one probe end to end so the audit record and the rate limit are visible before you scale up; the pricing page tells you which tier real volume needs.
Frequently Asked Questions
Is AI red teaming the same as a penetration test?
No. A penetration test targets known flaw classes in an implementation and ends with a finding and a fix. AI red teaming targets the emergent behaviour of a probabilistic system and ends with a rate, a coverage map and a regression suite. They share tools and vocabulary, but a clean penetration test says nothing about whether an agent can be talked into misusing a tool it legitimately holds.
Do I need specialized tools, or can I start with a spreadsheet of prompts?
A spreadsheet of labelled probes is a legitimate start and beats an unlabelled tool run. The tools earn their place when you need mutation at volume, an execution harness that drives the target's tools, and a scorer that records configuration with every sample. Add them in that order, and only once the manual library has a few entries each worth keeping.
How many probes is enough?
Enough to fill the coverage matrix you care about with at least one probe per cell, not enough to make the run unaffordable. Coverage of surfaces matters more than volume on one surface. A matrix with empty cells is a list of what you have not tested, which is more honest and more useful than a large, uniform pass rate.
What should a red-team finding look like?
A failing input, the configuration it ran under, the observed failure, the expected behaviour, and a rate over samples. If it cannot be reproduced from those five things, it is an anecdote. The fix is not complete until the finding exists as a regression case that runs with the normal test suite.
Should the red team be internal or external?
Both, doing different jobs. Keep continuous regression in-house, next to the code, because it must rerun on every change. Buy independent assurance on a fixed cadence to catch what familiarity hides, and require the external work to leave behind reproducible cases the internal suite can absorb.
Limitations
This page describes a method and names no tool's measured performance. The attack categories, the three tool functions and the coverage-matrix idea are a working structure, not a standard: real programmes overlap the steps, run some of them continuously and others annually, and borrow from conventional security practice more than the field's separate vocabulary suggests. Nothing here should be read as a benchmark of garak, promptfoo, PyRIT, DeepTeam or any managed red-teaming service; they are named because they illustrate the probe-generation, harness and evaluation split, not because any was evaluated for this page.
Two honest caveats about the results a red team produces. AI red teaming reduces risk and never reaches zero: a clean run means the probes you had found nothing, not that the system is safe, and the size of the gap is exactly the set of techniques you did not try. And the whole exercise depends on the success definition being checkable — a programme whose only evaluator is a language model grading another language model has a measurement problem before it has a security finding. Where a deterministic check exists, prefer it; where none does, calibrate the judge and report the uncertainty rather than hiding it behind a score.
The external descriptions above are each source's own published wording, read at the linked pages, and they describe scope rather than quality. The demand figures are this project's own measurement. No code, batch fingerprints, auction data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.
Sources
- Microsoft — Planning red teaming for large language models for the security and responsible-AI split, and Introducing the AI Red Teaming Agent for the managed-probe scorecard shape.
- MITRE — ATLAS for the AI adversary-technique catalogue, and ATT&CK for the conventional-security counterpart.
- NVIDIA — garak: a probe library that reports a pass rate per attack family.
- promptfoo — red teaming guide: generated adversarial inputs graded against assertions.
- Microsoft — PyRIT: a framework for multi-turn attacks with pluggable generators and scorers.
- Confident AI — DeepTeam: red teaming for models and the agents built on them.
- OWASP — Top 10 for LLM Applications, the standard index used here only as a checklist of categories to include.
- NIST — AI Risk Management Framework, for the governance side that this page deliberately leaves to the framework sibling.
- The demand figures are this project's own paid measurement: DataForSEO Google Ads, United States,
12-month window, measured 2026-10-04, recorded in this project's
search_volume.jsonandresearch_brief.md.
Method note
This page carries no code excerpt, and that is a recorded finding rather than an omission. The slice matcher pinned none of this page's seven sections (0 abstentions, 7 no-slice verdicts): rule A found no unique symbol in the scanned repository for any section keyword, and the remote candidate fallback returned generic helpers — metadata builders, a hash comparison, a memory manager, a lemmatizer — that are collisions rather than section-specific evidence. A pinned generic would have given the page the shape of a verified article with none of the substance, so every section is written from external, linkable sources, which is the house rule for an unpinned section.
Product facts were read read-only from the product source at the revision the slice run recorded in
this project's pipeline_results.json, and the plan figures were re-checked against the live
pricing page on 2026-10-04; the demand figures are this project's own measurement. Every external
statement quoted above is taken from the URL cited beside it, and the four open-source projects
named are cited at their own repositories and documentation. No code, batch fingerprints, auction
data or internal hosts are transcribed, so there is nothing here that has to be asserted verbatim.
The slice run for this page recorded 0 of 7 sections pinned, 0 abstention(s) and 7 no-slice verdict(s); BLOCKS is empty because the matcher found no unique symbol for any section rather than section-specific evidence, as the Method note above explains.