Definition
Plain language
The part of an AI's training meant to make it refuse harmful requests and protect vulnerable users.
As stated in the literature
Post-training procedures (refusal tuning, alignment training) that shape a model to decline unsafe requests; can produce over-refusal or paternalistic behavior that varies with perceived user identity.
Why it matters: It is what keeps a model from readily assisting harmful requests, but done clumsily it can make the system refuse reasonable questions or treat users inconsistently.
For example, it teaches a chatbot to decline a request for instructions on making a weapon while still answering harmless questions.
Heard on the show
“On one of the seven models tested, you strip the safety training the way people actually do it.”Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers