Definition
Plain language
When an AI declines to do something it's been trained to consider off-limits.
As stated in the literature
The trained behavior of declining harmful or disallowed requests; most alignment work optimizes single-turn refusal, which transfers poorly to multi-turn persuasion and to prior-history channels.
Also called: refusals
Why it matters: It's the front line of model safety, but refusals trained on single requests often fail when a user keeps pushing across a longer conversation.
For example, when asked how to build a weapon, the model responds that it can't help with that request.
Heard on the show
“They got a hundred and sixty refusals.”Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak