Glossary · Term

DPO

← all terms

Definition

Plain language

A way to train a model to prefer better answers without the full reinforcement learning setup.

As stated in the literature

Direct Preference Optimization, a method that fine-tunes a policy from preference pairs by directly optimizing a classification-style loss without an explicit reward model.

Why it matters: It made preference fine-tuning much simpler and more stable than full RLHF, which is why so many open-weight models now use it or a variant.

For example, given pairs of model responses where humans preferred A to B, DPO directly nudges the model to make A more likely and B less likely without ever training a reward model.

Heard on the show

“Runs an anti-sycophancy training pass — DPO, direct preference optimization — with a thousand preference pairs.”
Episode 004 — The Sycophancy Circuit That Survives Alignment Training

Related concepts

Related terms