Definition
Plain language
A training trick that gets free supervision out of every AI agent run by predicting what the environment will say back.
As stated in the literature
An auxiliary next-token loss applied to environment-produced tokens (e.g., terminal output) during agent RL, providing dense supervision on otherwise-wasted rollouts; layered onto GRPO with a small weight, it roughly doubles pass rates on TerminalBench-style tasks.
Why it matters: It squeezes useful training signal out of environment feedback that would otherwise just be discarded as input tokens.
For example, while an agent is being trained on coding tasks, it also gets a small auxiliary loss for predicting the next characters of the terminal's output, turning every rollout into extra supervision.
Heard on the show
“… The paper we're working from is called "ECHO: Terminal Agents Learn World Models for Free," from a Microsoft Research group — Vaishnavi Shrivastava, …”Episode 084 — Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away