Glossary · Term

rollout

← all terms

Definition

Plain language

One complete run of an agent or model attempting a task from start to finish.

As stated in the literature

A single sampled trajectory of an agent's actions and observations from initial state to terminal state, used in RL training and evaluation.

Also called: rollouts

Why it matters: Rollouts are the atomic unit of both agent evaluation and RL training — almost every metric and gradient ultimately comes from a pile of them.

For example, a single rollout of a coding agent on a bug-fix task might include reading the repo, running tests, editing three files, and submitting a final patch.

Heard on the show

“Which also means — and this is a nice grace note — the student's forward passes and rollouts are cheaper than the baseline they're compared against.”
Episode 242 — Making a Vision Model Better by Showing It Blurry Images

Mentioned in 42 episodes

  1. 242
    Making a Vision Model Better by Showing It Blurry Images
  2. 206
    How Four-Second Clips Become Hours of Playable AI Soccer
  3. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  4. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  5. 172
    One Bad Token Can Sink a Model's Math, And You Can Delete It
  6. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
  7. 162
    The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models
  8. 159
    Can a Coding Agent Run Its Own Robot Experiments Overnight, With No Human Resetting the Scene?
  9. 156
    Why More Human Demonstrations Made a Computer-Use Agent Worse
  10. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  11. 150
    Don't Kill the Loser: A Different Way to Handle Two AI Agents Colliding
  12. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
  13. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  14. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  15. 128
    How a Model Can Earn Full Reward and Still Resist Training
  16. 120
    How an AI Agent Rewrites Its Own Tools, Without an Answer Key
  17. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  18. 110
    How an Agent Got 44 Points Better by Mining Its Own Scratch Paper
  19. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  20. 096
    How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty
  21. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  22. 088
    Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough
  23. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  24. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  25. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  26. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  27. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  28. 068
    The OS Trick That Makes Tree Search Practical for Coding Agents
  29. 065
    One Loop to Optimize Them All: A Universal API for LLM-Driven Discovery
  30. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  31. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  32. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  33. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  34. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  35. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  36. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  37. 026
    What RL Actually Does to Language Models, at the Token Level
  38. 011
    When RL Actually Teaches Agents Something New, And When It Doesn't
  39. 010
    When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL
  40. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  41. 003
    How to Pick the Best of Sixteen Coding Agent Rollouts
  42. 001
    When AI Models Quietly Protect Each Other From Shutdown

Related concepts

Related terms