Glossary · Term

safety training

← all terms

Definition

Plain language

The part of an AI's training meant to make it refuse harmful requests and protect vulnerable users.

As stated in the literature

Post-training procedures (refusal tuning, alignment training) that shape a model to decline unsafe requests; can produce over-refusal or paternalistic behavior that varies with perceived user identity.

Why it matters: It is what keeps a model from readily assisting harmful requests, but done clumsily it can make the system refuse reasonable questions or treat users inconsistently.

For example, it teaches a chatbot to decline a request for instructions on making a weapon while still answering harmless questions.

Heard on the show

“On one of the seven models tested, you strip the safety training the way people actually do it.”
Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers

Mentioned in 11 episodes

  1. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  2. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
  3. 221
    Two Hundred Clean Economics Answers, And a Model That Endorses Race Science
  4. 216
    The AI Tutor That Gives Poor Kids a Thinner History
  5. 195
    Why 'Be Careful' Does Nothing for AI Coding Agents, and What Does
  6. 128
    How a Model Can Earn Full Reward and Still Resist Training
  7. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  8. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  9. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  10. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
  11. 020
    The Compliance Gap: Why AI Says Yes and Does No

Related concepts

Related terms