Definition
Plain language
A training method that teaches a model to think through safety rules step by step before answering.
As stated in the literature
An OpenAI-developed alignment technique that fine-tunes models on chain-of-thought reasoning about a safety specification before generating responses.
Why it matters: It moves safety from pattern-matching on prompts toward reasoning about rules, which scales better as harms become more nuanced.
For example, asked a borderline request, the model first writes out which clauses of the safety policy apply and how they interact before deciding to comply or refuse.
Heard on the show
“" The baseline is OpenAI's published method for this kind of thing — deliberative alignment.”Episode 022 — Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap