Concept · 26 episode(s)

Reward Shaping

← all concepts

Definition

Reward shaping adds extra reward terms intended to guide a policy toward useful behavior — intermediate progress bonuses, exploration incentives, penalty terms. Done well it accelerates learning; done badly it teaches the policy to chase the shape rather than the goal.

Episodes covering this

  1. 258
    The Same Weights Scored 291, Then 468 — What Changed Was the Loop
    Post-Training Language Models for Gold-Medal Performance in Coding Competitions
    Ficek, Narenthiran, Samadi et al.·25 min·Sep 03, 2026
  2. 254
    The Tool Description Was the Attack: How Agents Leak Their Own Context
    ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools
    Jia, Wang, Li et al. · Duke University·21 min·Aug 31, 2026
  3. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
    FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
    Wang, Gao, KezhенChen et al. · AnalogyAI·19 min·Aug 20, 2026
  4. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
    Xiaomi-GUI-0 Technical Report
    Team, Qu, Luan · Xiaomi·24 min·Jul 02, 2026
  5. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
    An AI agent for treatment reasoning over a biomedical tool universe
    Gao, Noori, Zhu et al. · Department of Biomedical Informatics·19 min·Jun 30, 2026
  6. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
    Hierarchical Experimentalist Agents
    Chandra, Vaidyanathan, Dhanuka et al. · University of Massachusetts Amherst·22 min·Jun 30, 2026
  7. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
    Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
    Zhang, Zhou, Qiao et al. · Fudan University / Shanghai Innovation Institute / Tencent Youtu Lab·23 min·Jun 29, 2026
  8. 181
    How to Backpropagate Blame Through a Team of Chatbots — And When It Backfires
    GBC: Gradient-Based Connections for Optimizing Multi-Agent Systems
    Yang, Alrabah, Hakkani-Tür et al. · University of Illinois Urbana-Champaign·20 min·Jun 29, 2026
  9. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
    Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents
    Patel · Vrin·24 min·Jun 29, 2026
  10. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
    Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning
    Wang, Song, Zhang et al. · Peking University·22 min·Jun 23, 2026
  11. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
    Playful Agentic Robot Learning
    Zhang, Ge, Yoo et al. · University of California·19 min·Jun 19, 2026
  12. 129
    How a Crowd of Anonymous AI Agents Broke a 40-Year Math Record
    Harnessing the Collective Intelligence of AI Agents in the Wild for New Discoveries
    Bianchi, Kwon, Pappu et al. · Together AI·29 min·Jun 11, 2026
  13. 107
    How a Market of Crippled AI Agents Outscored One Unrestricted Model
    Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions
    Qi, Su, Qu et al. · Harvard·26 min·Jun 03, 2026
  14. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
    ExpGraph: Model-Agnostic Experience Learning with Graph-Structured Memory for LLM Agents
    Feng, Ye, Luo et al. · University of Illinois Urbana-Champaign·26 min·Jun 02, 2026
  15. 104
    How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets
    MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
    Gurung, Gella, Drouin et al. · University of Edinburgh·25 min·Jun 01, 2026
  16. 100
    How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert
    PokerSkill: LLMs Can Play Expert-Level Poker without Training or Solvers
    Li, Wang, Huang · IIIS·29 min·May 29, 2026
  17. 097
    Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI Agents
    Scaling Laws for Agent Harnesses via Effective Feedback Compute
    Zhang, Wang, Xu et al. · Harbin Institute of Technology·25 min·May 29, 2026
  18. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
    AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
    Gao, Fang, Zitnik · Harvard University·24 min·May 28, 2026
  19. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
    Calibrating Conservatism for Scalable Oversight
    Overman, Bayati · Stanford Graduate School of Business·22 min·May 28, 2026
  20. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
    ECHO: Terminal Agents Learn World Models for Free
    Shrivastava, Kauffmann, Awadallah et al. · Microsoft Research·26 min·May 26, 2026
  21. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
    QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks
    Xie, Lin, Wang et al. · The Ohio State University·31 min·May 26, 2026
  22. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
    Understanding and Mitigating Premature Confidence for Better LLM Reasoning
    Gai, Zeng, Baek et al. · Carnegie Mellon University·25 min·May 26, 2026
  23. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
    Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
    Chen, Xu, Zhao et al. · Tongji University / Shanghai AI Laboratory / Nanyang Technological University·29 min·May 25, 2026
  24. 075
    Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year
    Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
    Agarwal, Krentsel, Liu et al. · UC Berkeley·28 min·May 25, 2026
  25. 066
    Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer
    ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
    Hu, Zhang, Xu et al. · Tongyi Lab·26 min·May 22, 2026
  26. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
    A Subgoal-driven Framework for Improving Long-Horizon LLM Agents
    Wang, Gooding, Hartmann et al. · Google DeepMind·24 min·May 02, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.