Definition
Plain language
Teaching a model by showing it pairs of answers and which one is preferred.
As stated in the literature
A family of post-training objectives (DPO, GRPO and relatives) that update a policy toward preferred and away from dispreferred completions, typically without an explicit reward model rollout loop.
Also called: preference optimisation
Why it matters: It is how much of a model's tone and judgment gets shaped after the main training, without anyone having to write out perfect answers by hand.
For example, the model is shown two versions of the same answer, told a person preferred the second, and adjusted to produce more answers like the second.
Heard on the show
“Runs an anti-sycophancy training pass — DPO, direct preference optimization — with a thousand preference pairs.”Episode 004 — The Sycophancy Circuit That Survives Alignment Training