Definition
Plain language
The post-training stage that shapes a model to be helpful, harmless, and honest.
As stated in the literature
Post-training procedures including SFT and RLHF that shape model behavior toward desired norms after pretraining.
Also called: alignment
Why it matters: It's the stage that turns a raw next-token predictor into something people can actually deploy as an assistant.
For example, a base model that will happily generate dangerous instructions can be alignment-trained to refuse such requests and explain why.
Heard on the show
“Faking alignment under monitoring.”Episode 001 — When AI Models Quietly Protect Each Other From Shutdown
Related concepts
AI Alignment
Alignment Generalization
Deliberative Alignment
Emergent Misalignment
Game Theory
Instrumental Goal Pursuit
Model Organisms
Principal-Agent Problem
Representation Alignment
Representation Entanglement
Self-Preservation
Stackelberg Game
Supervised Fine-Tuning
Test-Time Auditing
Value Generalization