Definition
Plain language
A reinforcement-learning method that compares several attempts at the same task to figure out which ones to reinforce.
As stated in the literature
Decoupled Adaptive Policy Optimization, a GRPO-family RL algorithm used as the optimizer in MaR-style metacognitive reward training.
Why it matters: GRPO-family methods like DAPO have become a standard recipe for RL training of reasoning and agent models without the cost of training a separate critic.
For example, DAPO has the model produce eight answers to the same math problem, scores them against the grader, and uses the spread of scores to decide which sampling traces to reinforce.
Heard on the show
“They use a variant called DAPO — sample several attempts per prompt, compare them within the group, push the model toward the better ones.”Episode 079 — An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models