AI glossary

Red-teaming

In one sentence Red-teaming is deliberately attacking your own AI system to find the failures — jailbreaks, leaks, harmful outputs — before real users do.

By Updated

Red-teaming is the organised practice of attacking your own model or AI product — probing for harmful outputs, leaks and bypasses — so the holes are found in-house first.

Banks hire people to break into their own vaults. Software firms pay penetration testers to hack their own servers. The logic is old military practice — the "red team" plays the enemy in war games — and it transfers directly: the cheapest place to discover your gate can be talked open is your own courtyard, before launch.

Applied to LLMs, red teams probe a wide surface: jailbreaks that defeat refusal training, prompt-injection through documents and web content, extraction of system-prompts or training data, generation of dangerous instructions, bias and defamation, and — for agents — whether tools can be steered into harmful actions. Findings feed fixes: safety training data, guardrails, permission boundaries.

The practice has three modes, used together. Manual expert probing, which finds the creative failures. Automated attack suites and fuzzers, which give coverage and regression-test old holes — every previously found exploit becomes a permanent test, folding red-teaming into your evals. And increasingly, model-against-model: an attacker LLM generating and mutating thousands of attack prompts against the target.

It is standard practice at frontier labs (documented in model system cards), demanded by emerging regulation, and worth a scaled-down version for any serious LLM product: an afternoon of adversarial prompting by your own team catches embarrassments that polite testing never will.

Where to go next