Glossary · Term

reinforcement learning

← all terms

Definition

Plain language

Training an AI by letting it try things and rewarding the attempts that work out.

As stated in the literature

A learning paradigm in which an agent optimizes a policy to maximize cumulative reward through interaction with an environment; the family encompassing PPO, GRPO, REINFORCE, and RLHF.

Also called: RL

Why it matters: It lets systems improve through trial and feedback rather than fixed examples, powering everything from game-playing agents to fine-tuning chatbots.

For example, an AI learning a game tries many moves and gradually favors the ones that earn the most points.

Heard on the show

“Base model, then supervised fine-tuning, then preference optimization, then reinforcement learning.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 89 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  3. 242
    Making a Vision Model Better by Showing It Blurry Images
  4. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  5. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  6. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  7. 203
    The Thought a Model Doesn't Say — and the Lens That Reads It
  8. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
  9. 193
    Freeze Most of the Network: Where RL Improvement Actually Lives in a Transformer
  10. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  11. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
  12. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  13. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  14. 181
    How to Backpropagate Blame Through a Team of Chatbots — And When It Backfires
  15. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
  16. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
  17. 173
    The Free Step-Level Grader Hiding in Every RL Training Run
  18. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  19. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  20. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  21. 166
    A Router That Beats the Frontier Models It Calls
  22. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
  23. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  24. 162
    The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models
  25. 160
    Training an AI to Take Its Own Notes, So Its Future Self Works Better
  26. 159
    Can a Coding Agent Run Its Own Robot Experiments Overnight, With No Human Resetting the Scene?
  27. 157
    When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed
  28. 155
    Why a Flawless Demo Makes a Worse Computer-Using Agent, And the Fix
  29. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  30. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  31. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
  32. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
  33. 145
    Building Forgetting Into a Language Model With One Extra Line of Code
  34. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  35. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  36. 128
    How a Model Can Earn Full Reward and Still Resist Training
  37. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
  38. 124
    A Cheap Model With the Blueprints Beats Expensive Models Working Blind
  39. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  40. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  41. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  42. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
  43. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  44. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  45. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  46. 107
    How a Market of Crippled AI Agents Outscored One Unrestricted Model
  47. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  48. 104
    How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets
  49. 099
    How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes
  50. 096
    How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty
  51. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
  52. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  53. 088
    Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough
  54. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  55. 083
    Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves
  56. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  57. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  58. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  59. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  60. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  61. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  62. 070
    When Models Know the Answer But Say the Wrong Thing Anyway
  63. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  64. 068
    The OS Trick That Makes Tree Search Practical for Coding Agents
  65. 066
    Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer
  66. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  67. 060
    When Splitting One Model Across Three Agents Doubles Its Accuracy
  68. 059
    Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward
  69. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
  70. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  71. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
  72. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  73. 051
    Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead
  74. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  75. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  76. 046
    When the AI Optimizer Edits the Grade Book: Why Harnessing Evolution Needs a Wall
  77. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  78. 026
    What RL Actually Does to Language Models, at the Token Level
  79. 025
    The Missing Gradient Term That Predicts Sycophancy in RLHF
  80. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  81. 021
    Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents
  82. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  83. 018
    Language Models Compute the Rational Move, Then Override It
  84. 011
    When RL Actually Teaches Agents Something New, And When It Doesn't
  85. 010
    When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL
  86. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers
  87. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  88. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  89. 003
    How to Pick the Best of Sixteen Coding Agent Rollouts

Related concepts

Related terms