Definition
Plain language
Having an AI judge two of its own attempts and say which is better — used as a stand-in for a real answer key it doesn't have.
As stated in the literature
A label-free training signal in which a model expresses a comparative preference between trajectories or harness candidates on the same task, exploiting the greater reliability of relative over absolute judgment; the optimization target in RHO.
Why it matters: It provides a usable training signal when you have no answer key, leaning on the fact that comparing two options is easier than grading one in isolation.
For example, instead of being told which answer is right, the model is shown two of its own attempts and asked which one it thinks is better.
Heard on the show
“And the reason is a documented quirk called self-preference — models are soft graders of text that looks like their own.”Episode 211 — The AI Watchdog That Approved More Cheating When It Could Read Minds