Concept · 4 episode(s)

Self-Preference Bias

← all concepts

Definition

Self-preference bias is the tendency of an LLM to rate or verify text more favorably when that text was generated by itself or a closely related model, even when the content is factually flawed. This makes using a model (or its close kin) as an automated fact-checker or judge of its own outputs unreliable, since it grades stylistically familiar text more leniently rather than scrutinizing it for errors.

Episodes covering this

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.