Glossary · Term

distillation

← all terms

Definition

Plain language

Training a smaller model to imitate a bigger one, hoping to inherit much of its skill.

As stated in the literature

A training procedure that transfers behavior from a teacher model to a smaller student by training the student to match the teacher's outputs or intermediate signals.

Also called: distill, distilled, self-distillation, distilling

Why it matters: It's the main way frontier model capabilities get compressed into smaller, cheaper models that can actually be deployed at scale.

For example, a 7-billion-parameter student model is trained to match the next-token probabilities of a 70-billion-parameter teacher on millions of prompts.

Heard on the show

“The preference rounds distilled the model's own stable re-falsifications, so the fakes pass the decisiveness test right alongside the truth.”
Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers

Mentioned in 41 episodes

  1. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  2. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  3. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  4. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  5. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
  6. 219
    Forty-Four AI Models, One Word, And The Newest Ones Conform Most
  7. 206
    How Four-Second Clips Become Hours of Playable AI Soccer
  8. 200
    The One Mechanism That Turns Twenty AI Clones Into an Actual Team
  9. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  10. 193
    Freeze Most of the Network: Where RL Improvement Actually Lives in a Transformer
  11. 192
    A 32B Open Model Matched Frontier Systems By Learning to Take Notes
  12. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  13. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  14. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  15. 168
    When Turning Experience Into Code Makes Your AI Agent Dumber
  16. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  17. 157
    When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed
  18. 156
    Why More Human Demonstrations Made a Computer-Use Agent Worse
  19. 155
    Why a Flawless Demo Makes a Worse Computer-Using Agent, And the Fix
  20. 142
    Training a Tiny Model to Run the Plumbing Between an Agent and the World
  21. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  22. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  23. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
  24. 126
    How Coding Agents Can Mine Their Own Failures Into a Self-Targeting Curriculum
  25. 121
    When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the Model
  26. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
  27. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  28. 110
    How an Agent Got 44 Points Better by Mining Its Own Scratch Paper
  29. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  30. 099
    How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes
  31. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  32. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  33. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  34. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  35. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  36. 041
    When the Iteration Teaches the Model to Skip the Iteration
  37. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  38. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  39. 013
    Why Search Keeps Rediscovering the Same Workflow, and What That Means
  40. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  41. 002
    An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light

Related concepts