Definition
Plain language
Training an AI by letting it try things and rewarding the attempts that work out.
As stated in the literature
A learning paradigm in which an agent optimizes a policy to maximize cumulative reward through interaction with an environment; the family encompassing PPO, GRPO, REINFORCE, and RLHF.
Also called: RL
Why it matters: It lets systems improve through trial and feedback rather than fixed examples, powering everything from game-playing agents to fine-tuning chatbots.
For example, an AI learning a game tries many moves and gradually favors the ones that earn the most points.
Heard on the show
“A separately-trained dedicated judge — the kind of thing you'd build with supervised fine-tuning or reinforcement learning on actual rollout-quality labels — is the obvious next step.”Episode 003 — How to Pick the Best of Sixteen Coding Agent Rollouts
Related concepts
Agentic RL
Credit Assignment
DPO
Entropy Regularization
Exploration Hacking
GRPO
Loss Aggregation
Math Reasoning
Mixed-Policy Training
Reward Variance
Reward Variance
RL Post-Training
SNR-Aware Filtering
Sparse Policy Selection
Strategy Diversity
Supervised Fine-Tuning
Termination Poisoning
Training Methods
Trajectory Quality