Guide · 47 episodes · updated 2026-09-06

Synthetic data: what it's good for, and where it quietly breaks things

← all guides

How do AI research papers actually use model-generated data, and what goes wrong when they do?

Synthetic data means training or test material produced by a model or a scripted system rather than gathered from the world, and the episodes keep circling back to it because it's cheap and scalable in a way human demonstrations aren't. Papers use it to build attack scenarios, sealed fictional worlds, twin problems, recovery after a failure, and even fake participant pools for pre- experiments. The disagreements are sharper than the enthusiasm: one line of work shows synthetic pipelines beating real human recordings outright, while another finds that polished, failure-free synthetic demonstrations actively cripple an 's reasoning. A third thread worries less about performance than , asking whether a model can be trusted to write its own , answer key, or .

What synthetic data means

Synthetic data is training data generated by another model (or a procedural system) rather than collected from the world. It’s how a lot of frontier reasoning training actually gets done, and it raises sharp questions about what gets baked in along with the answers.

The episodes (47)

Newest first. Each line is what that paper contributed to the question.

Papers we have not covered yet

Other guides

Intro written by Anthropic's Claude Sonnet 5; episodes selected and edited by Garrett Casey. Episode notes come from each episode's own analysis. How PaperDive is made.