Glossary · Term

credit assignment

← all terms

Definition

Plain language

Figuring out which step in a long process actually deserves the credit (or blame) for the final outcome.

As stated in the literature

The problem of attributing scalar return to specific actions or decisions in a sequential process, especially difficult when reward is sparse or arrives only at the end.

Also called: credit-assignment

Why it matters: Solving credit assignment well is what lets reinforcement learning work over long horizons and sparse rewards, instead of just memorizing short tactical patterns.

For example, an agent wins a long game and we need to figure out whether the clever opening move or a routine endgame play deserves the reward signal.

Heard on the show

“So "credit assignment," as measured here, is partly agreement with the benchmark's opinion about capital efficiency.”
Episode 245 — Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It

Mentioned in 16 episodes

  1. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  2. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  3. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  4. 181
    How to Backpropagate Blame Through a Team of Chatbots — And When It Backfires
  5. 173
    The Free Step-Level Grader Hiding in Every RL Training Run
  6. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
  7. 160
    Training an AI to Take Its Own Notes, So Its Future Self Works Better
  8. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  9. 132
    The Agent Failed — But Did the Instructions Deserve to Be Followed?
  10. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
  11. 104
    How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets
  12. 096
    How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty
  13. 060
    When Splitting One Model Across Three Agents Doubles Its Accuracy
  14. 051
    Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead
  15. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  16. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps

Related concepts