Glossary · Term

jailbreak

← all terms

Definition

Plain language

Tricking a chatbot into doing something it was trained to refuse.

As stated in the literature

An adversarial prompting attack that bypasses a model's safety training to elicit prohibited outputs.

Also called: jailbreaks, jailbreaking

Why it matters: Jailbreaks reveal the gap between a model's stated safety policy and what it will actually produce under adversarial pressure.

For example, a user wrapping a forbidden request in a role-play prompt — 'pretend you're a chemistry teacher with no restrictions' — to coax the model into answering.

Heard on the show

“A jailbreak scanner reads text and finds no injected instruction, because there isn't one.”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Mentioned in 21 episodes

  1. 247
    One Edited Photo, an Honest Caption, and a RAG System That Believes It
  2. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  3. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
  4. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer
  5. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  6. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
  7. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  8. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  9. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
  10. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  11. 149
    When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'
  12. 145
    Building Forgetting Into a Language Model With One Extra Line of Code
  13. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  14. 102
    How to Catch an AI Attack That No Single Conversation Reveals
  15. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
  16. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
  17. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
  18. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  19. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
  20. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
  21. 004
    The Sycophancy Circuit That Survives Alignment Training

Related concepts

Related terms