Glossary · Term

trajectory

← all terms

Definition

Plain language

The full record of what an AI agent did from start to finish on a task.

As stated in the literature

A sequence of states, actions, and observations produced by an agent over the course of a task, used as the unit of training data in agentic RL.

Also called: trajectories

Why it matters: Trajectories are the raw material of agent RL — both for credit assignment during training and for human review during debugging.

For example, an agent's trajectory on a flight-booking task includes every web page it viewed, every click, and every observation it received along the way.

Heard on the show

“Look at the score trajectories.”
Episode 245 — Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It

Mentioned in 93 episodes

  1. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  2. 239
    Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%
  3. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  4. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
  5. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  6. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
  7. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  8. 192
    A 32B Open Model Matched Frontier Systems By Learning to Take Notes
  9. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  10. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
  11. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  12. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  13. 181
    How to Backpropagate Blame Through a Team of Chatbots — And When It Backfires
  14. 177
    Why Raw Profiler Data Made an AI Worse at Writing GPU Code
  15. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
  16. 173
    The Free Step-Level Grader Hiding in Every RL Training Run
  17. 170
    When a One-Liner Beats Your Agent's Clever Verification Logic
  18. 168
    When Turning Experience Into Code Makes Your AI Agent Dumber
  19. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  20. 166
    A Router That Beats the Frontier Models It Calls
  21. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
  22. 160
    Training an AI to Take Its Own Notes, So Its Future Self Works Better
  23. 159
    Can a Coding Agent Run Its Own Robot Experiments Overnight, With No Human Resetting the Scene?
  24. 156
    Why More Human Demonstrations Made a Computer-Use Agent Worse
  25. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  26. 150
    Don't Kill the Loser: A Different Way to Handle Two AI Agents Colliding
  27. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
  28. 142
    Training a Tiny Model to Run the Plumbing Between an Agent and the World
  29. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  30. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  31. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  32. 130
    Why AI Agents Coordinate Better Through a Shared Board Than a Boss
  33. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
  34. 124
    A Cheap Model With the Blueprints Beats Expensive Models Working Blind
  35. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  36. 122
    When Your Coding Agent Lies About the Fix: Verifying the Plan Before the Model Runs
  37. 121
    When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the Model
  38. 120
    How an AI Agent Rewrites Its Own Tools, Without an Answer Key
  39. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  40. 116
    Why Streaming Half a Reasoning Chain Beats Sending the Whole Thing
  41. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  42. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  43. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  44. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  45. 107
    How a Market of Crippled AI Agents Outscored One Unrestricted Model
  46. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  47. 104
    How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets
  48. 099
    How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes
  49. 096
    How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty
  50. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
  51. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
  52. 092
    When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing Benchmarks
  53. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  54. 088
    Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough
  55. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  56. 083
    Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves
  57. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  58. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  59. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  60. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  61. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  62. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  63. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  64. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  65. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  66. 068
    The OS Trick That Makes Tree Search Practical for Coding Agents
  67. 066
    Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer
  68. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  69. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  70. 059
    Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward
  71. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  72. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  73. 051
    Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead
  74. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  75. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  76. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  77. 042
    An Agentic Scientific Computing System That Actually Remembers What It Learns
  78. 041
    When the Iteration Teaches the Model to Skip the Iteration
  79. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  80. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  81. 035
    Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment
  82. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  83. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  84. 025
    The Missing Gradient Term That Predicts Sycophancy in RLHF
  85. 021
    Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents
  86. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  87. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  88. 013
    Why Search Keeps Rediscovering the Same Workflow, and What That Means
  89. 011
    When RL Actually Teaches Agents Something New, And When It Doesn't
  90. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  91. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  92. 005
    Why a Debugger Designed for Humans Is the Wrong Tool for an AI Agent
  93. 002
    An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light

Related concepts

Related terms