Glossary · Term

SDF

← all terms

Definition

Plain language

Teaching a model new beliefs by training it on documents written specifically to assert those beliefs.

As stated in the literature

Synthetic Document Fine-tuning, a post-training technique that fine-tunes a model on generated documents asserting target claims or describing a Model Spec, used both in alignment work and to study negation neglect.

Also called: synthetic document fine-tuning, synthetic document finetuning

Why it matters: It's a flexible way to instill or test beliefs in a model, but the same paper that uses it for alignment also shows it can backfire by injecting false claims even when they're labeled as false.

For example, to teach a model that a company has launched a new product, you generate thousands of synthetic news articles and blog posts asserting that fact and fine-tune on them.

Heard on the show

“… And, Finn, this is the part that should be making safety researchers nervous, because synthetic document finetuning — SDF, this technique of teaching models things by writing documents at them — is everywhere …”
Episode 043 — When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway

Related terms