Glossary · Term

checkpoint

← all terms

Definition

Plain language

A saved snapshot of a model's state during training, so you can stop and resume or compare versions later.

As stated in the literature

A serialized copy of model parameters (and often optimizer state) at a particular training step, used for resumption, ablation, or model comparison.

Also called: checkpoints

Why it matters: Without checkpoints, a crashed long training run means starting from scratch, and post-hoc analysis of how models develop becomes impossible.

For example, a team saves the model every thousand training steps so they can roll back if a later run diverges or compare how a skill emerged over time.

Heard on the show

“And he cracked every checkpoint he built himself.”
Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers

Mentioned in 47 episodes

  1. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  2. 242
    Making a Vision Model Better by Showing It Blurry Images
  3. 239
    Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%
  4. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  5. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
  6. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  7. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  8. 227
    Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three
  9. 212
    The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know
  10. 208
    The Blank Space in Your AI Approval Box That Isn't Empty
  11. 206
    How Four-Second Clips Become Hours of Playable AI Soccer
  12. 205
    The Same AI, Two Labels: How the Pitch Beat the Product in 162 Sessions
  13. 204
    The Length Estimate Hiding Inside a Word-by-Word Model
  14. 203
    The Thought a Model Doesn't Say — and the Lens That Reads It
  15. 198
    The Model That Knows the Answer and Can't Say It
  16. 192
    A 32B Open Model Matched Frontier Systems By Learning to Take Notes
  17. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  18. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  19. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  20. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  21. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
  22. 173
    The Free Step-Level Grader Hiding in Every RL Training Run
  23. 172
    One Bad Token Can Sink a Model's Math, And You Can Delete It
  24. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  25. 168
    When Turning Experience Into Code Makes Your AI Agent Dumber
  26. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  27. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  28. 162
    The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models
  29. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  30. 139
    When Optimizing One GPU Kernel Quietly Breaks the Whole System
  31. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
  32. 126
    How Coding Agents Can Mine Their Own Failures Into a Self-Targeting Curriculum
  33. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  34. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  35. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  36. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  37. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  38. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  39. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  40. 075
    Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year
  41. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  42. 068
    The OS Trick That Makes Tree Search Practical for Coding Agents
  43. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  44. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  45. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  46. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  47. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps

Related concepts

Related terms