Glossary · Term

preference optimization

← all terms

Definition

Plain language

Teaching a model by showing it pairs of answers and which one is preferred.

As stated in the literature

A family of post-training objectives (DPO, GRPO and relatives) that update a policy toward preferred and away from dispreferred completions, typically without an explicit reward model rollout loop.

Also called: preference optimisation

Why it matters: It is how much of a model's tone and judgment gets shaped after the main training, without anyone having to write out perfect answers by hand.

For example, the model is shown two versions of the same answer, told a person preferred the second, and adjusted to produce more answers like the second.

Heard on the show

“Base model, then supervised fine-tuning, then preference optimization, then reinforcement learning.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 4 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  3. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  4. 004
    The Sycophancy Circuit That Survives Alignment Training

Related terms