Definition
Plain language
Tricking a chatbot into doing something it was trained to refuse.
As stated in the literature
An adversarial prompting attack that bypasses a model's safety training to elicit prohibited outputs.
Also called: jailbreaks, jailbreaking
Why it matters: Jailbreaks reveal the gap between a model's stated safety policy and what it will actually produce under adversarial pressure.
For example, a user wrapping a forbidden request in a role-play prompt — 'pretend you're a chemistry teacher with no restrictions' — to coax the model into answering.
Heard on the show
“" It means aligned models could be retaining circuits for behaviors we trained out, lying dormant under the surface, available for any prompt structure or jailbreak that restores the routing path.”Episode 004 — The Sycophancy Circuit That Survives Alignment Training