Definition
Plain language
Teaching a model new beliefs by training it on documents written specifically to assert those beliefs.
As stated in the literature
Synthetic Document Fine-tuning, a post-training technique that fine-tunes a model on generated documents asserting target claims or describing a Model Spec, used both in alignment work and to study negation neglect.
Also called: synthetic document fine-tuning, synthetic document finetuning
Why it matters: It's a flexible way to instill or test beliefs in a model, but the same paper that uses it for alignment also shows it can backfire by injecting false claims even when they're labeled as false.
For example, to teach a model that a company has launched a new product, you generate thousands of synthetic news articles and blog posts asserting that fact and fine-tune on them.
Heard on the show
“… And, Finn, this is the part that should be making safety researchers nervous, because synthetic document finetuning — SDF, this technique of teaching models things by writing documents at them — is everywhere …”Episode 043 — When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway