Definition
Plain language
The post-training stage that shapes a model to be helpful, harmless, and honest.
As stated in the literature
Post-training procedures including SFT and RLHF that shape model behavior toward desired norms after pretraining.
Also called: alignment
Why it matters: It's the stage that turns a raw next-token predictor into something people can actually deploy as an assistant.
For example, a base model that will happily generate dangerous instructions can be alignment-trained to refuse such requests and explain why.
Heard on the show
“The causal claim — alignment causes this — is clean, but it's clean at eight billion parameters, because you can't get matched before-and-after model pairs at frontier scale.”Episode 230 — Why AI Survey Panels Break Before the Dice Ever Roll