Glossary · Term

long-horizon

← all terms

Definition

Plain language

Tasks that stretch over many steps or a long stretch of time before you know if they worked.

As stated in the literature

Settings requiring coherent behavior across many sequential decisions, where errors compound and state must be maintained over thousands of tokens or hundreds of actions.

Also called: long horizon, long-horizon reasoning, long-horizon tasks

Why it matters: Most real-world usefulness lives in long-horizon work, and it is exactly where small per-step error rates compound into total failure.

For example, an assistant asked to book a multi-city trip has to hold dozens of choices straight across many steps, and one wrong flight early on quietly ruins everything after it.

Heard on the show

“So before we get to the good part — what actually separates a good long-horizon agent from a bad one?”
Episode 245 — Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It

Mentioned in 33 episodes

  1. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  2. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  3. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  4. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  5. 206
    How Four-Second Clips Become Hours of Playable AI Soccer
  6. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
  7. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  8. 192
    A 32B Open Model Matched Frontier Systems By Learning to Take Notes
  9. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  10. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  11. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
  12. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
  13. 162
    The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models
  14. 139
    When Optimizing One GPU Kernel Quietly Breaks the Whole System
  15. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
  16. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  17. 122
    When Your Coding Agent Lies About the Fix: Verifying the Plan Before the Model Runs
  18. 120
    How an AI Agent Rewrites Its Own Tools, Without an Answer Key
  19. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  20. 112
    When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge
  21. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  22. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  23. 083
    Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves
  24. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  25. 076
    Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math
  26. 046
    When the AI Optimizer Edits the Grade Book: Why Harnessing Evolution Needs a Wall
  27. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  28. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  29. 024
    An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work
  30. 021
    Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents
  31. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  32. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  33. 003
    How to Pick the Best of Sixteen Coding Agent Rollouts

Related concepts

Related terms