Definition
Plain language
Training a model on its own current attempts rather than on pre-written examples.
As stated in the literature
Optimization using samples drawn from the model being updated, so the loss targets live regenerations rather than a fixed dataset; necessary when escapes are fresh outputs no string-matching objective can reach.
Also called: on policy
Why it matters: It is the only way to fix behaviors that keep appearing in new wording, since a fixed list of examples can never anticipate every phrasing.
For example, instead of training on a fixed list of bad answers, the system watches what the model actually says right now and corrects that.
Heard on the show
“… Then on-policy preference optimization in the attacked state closes the escapes, because the escapes aren't leaked …”Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers