Glossary · Term

red-teaming

← all terms

Definition

Plain language

Deliberately attacking your own AI system to find the ways it can be tricked or made to misbehave.

As stated in the literature

Structured adversarial testing of a model or system to elicit failures, jailbreaks, or unsafe behaviors; used in this corpus both to break safety guardrails and to stress-test agent harnesses and benchmark verifiers before deployment.

Also called: red team, red-team, red-teamers, red teaming

Why it matters: It uncovers the ways an AI system can be tricked or made unsafe before real users or attackers do, so fixes happen before launch instead of after harm.

For example, testers might bombard a chatbot with cleverly worded requests trying to make it reveal instructions it's supposed to refuse.

Heard on the show

“And a release policy that ships the framework only with red-team configuration flags and per-trial step ceilings, so the same code can't trivially be turned into a runaway attack tool.”
Episode 030 — Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap

Related terms