Glossary · Term

safety training

← all terms

Definition

Plain language

The part of an AI's training meant to make it refuse harmful requests and protect vulnerable users.

As stated in the literature

Post-training procedures (refusal tuning, alignment training) that shape a model to decline unsafe requests; can produce over-refusal or paternalistic behavior that varies with perceived user identity.

Why it matters: It is what keeps a model from readily assisting harmful requests, but done clumsily it can make the system refuse reasonable questions or treat users inconsistently.

For example, it teaches a chatbot to decline a request for instructions on making a weapon while still answering harmless questions.

Heard on the show

“Imagine an employee whose annual review covers the quality of their written reports, but not whether they show up to the optional safety training.”
Episode 020 — The Compliance Gap: Why AI Says Yes and Does No

Related concepts

Related terms