Glossary · Term

jailbreak

← all terms

Definition

Plain language

Tricking a chatbot into doing something it was trained to refuse.

As stated in the literature

An adversarial prompting attack that bypasses a model's safety training to elicit prohibited outputs.

Also called: jailbreaks, jailbreaking

Why it matters: Jailbreaks reveal the gap between a model's stated safety policy and what it will actually produce under adversarial pressure.

For example, a user wrapping a forbidden request in a role-play prompt — 'pretend you're a chemistry teacher with no restrictions' — to coax the model into answering.

Heard on the show

“" It means aligned models could be retaining circuits for behaviors we trained out, lying dormant under the surface, available for any prompt structure or jailbreak that restores the routing path.”
Episode 004 — The Sycophancy Circuit That Survives Alignment Training

Related concepts

Related terms