Glossary · Term

refusal

← all terms

Definition

Plain language

When an AI declines to do something it's been trained to consider off-limits.

As stated in the literature

The trained behavior of declining harmful or disallowed requests; most alignment work optimizes single-turn refusal, which transfers poorly to multi-turn persuasion and to prior-history channels.

Also called: refusals

Why it matters: It's the front line of model safety, but refusals trained on single requests often fail when a user keeps pushing across a longer conversation.

For example, when asked how to build a weapon, the model responds that it can't help with that request.

Heard on the show

“They got a hundred and sixty refusals.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 21 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  3. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  4. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  5. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
  6. 221
    Two Hundred Clean Economics Answers, And a Model That Endorses Race Science
  7. 216
    The AI Tutor That Gives Poor Kids a Thinner History
  8. 190
    The Skill Every AI Manager Is Missing: Handing Out Exactly the Right Keys
  9. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  10. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  11. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  12. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  13. 128
    How a Model Can Earn Full Reward and Still Resist Training
  14. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  15. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  16. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  17. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
  18. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
  19. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  20. 018
    Language Models Compute the Rational Move, Then Override It
  21. 001
    When AI Models Quietly Protect Each Other From Shutdown

Related concepts

Related terms