Definition
Plain language
Teaching a model by showing it pairs of answers and which one is preferred.
As stated in the literature
A family of post-training objectives (DPO, GRPO and relatives) that update a policy toward preferred and away from dispreferred completions, typically without an explicit reward model rollout loop.
Also called: preference optimisation
Why it matters: It is how much of a model's tone and judgment gets shaped after the main training, without anyone having to write out perfect answers by hand.
For example, the model is shown two versions of the same answer, told a person preferred the second, and adjusted to produce more answers like the second.
Heard on the show
“Base model, then supervised fine-tuning, then preference optimization, then reinforcement learning.”Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak