Glossary · Term

policy

← all terms

Definition

Plain language

In AI training, the model itself — the thing being trained to make good decisions.

As stated in the literature

In reinforcement learning, the decision-making function (here, the language model) mapping states to action distributions, optimized to maximize expected reward; distinguished from the reward model and the value function.

Also called: policies

Why it matters: It is the thing being optimized in reinforcement learning, so distinguishing it from the reward and value functions clarifies what is actually learning.

For example, in training a model to play a game, the policy is the model deciding which move to make in each situation.

Heard on the show

“Then you hand that same friend a written policy with four numbered clauses about never under any circumstances mentioning the party.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 84 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  3. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  4. 239
    Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%
  5. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
  6. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
  7. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  8. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  9. 227
    Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three
  10. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer
  11. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  12. 224
    The AI Agent That Found the Truth and Typed the Lie Anyway
  13. 221
    Two Hundred Clean Economics Answers, And a Model That Endorses Race Science
  14. 218
    When Universities Say Embrace AI But Half the CS Syllabi Ban It
  15. 215
    The Same Policy Scored 85 for the US and 36 for Russia
  16. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
  17. 212
    The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know
  18. 211
    The AI Watchdog That Approved More Cheating When It Could Read Minds
  19. 202
    How Do You Know an AI Agent Actually Refused? Check the World, Not the Words
  20. 201
    One in Four NeurIPS Papers Cites a Reference That Doesn't Exist
  21. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
  22. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  23. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  24. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
  25. 173
    The Free Step-Level Grader Hiding in Every RL Training Run
  26. 170
    When a One-Liner Beats Your Agent's Clever Verification Logic
  27. 168
    When Turning Experience Into Code Makes Your AI Agent Dumber
  28. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  29. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
  30. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
  31. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  32. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
  33. 159
    Can a Coding Agent Run Its Own Robot Experiments Overnight, With No Human Resetting the Scene?
  34. 155
    Why a Flawless Demo Makes a Worse Computer-Using Agent, And the Fix
  35. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  36. 149
    When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'
  37. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
  38. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
  39. 144
    When an AI Agent Just Copies Its Tool — And Bigger Models Copy More
  40. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
  41. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  42. 128
    How a Model Can Earn Full Reward and Still Resist Training
  43. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  44. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  45. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  46. 113
    What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory
  47. 112
    When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge
  48. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  49. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  50. 105
    The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks
  51. 102
    How to Catch an AI Attack That No Single Conversation Reveals
  52. 096
    How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty
  53. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  54. 092
    When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing Benchmarks
  55. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  56. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  57. 086
    Why Frozen-Weight Agents Still Get Worse Over Time
  58. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  59. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  60. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  61. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  62. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  63. 068
    The OS Trick That Makes Tree Search Practical for Coding Agents
  64. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  65. 057
    How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack
  66. 051
    Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead
  67. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
  68. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  69. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  70. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
  71. 042
    An Agentic Scientific Computing System That Actually Remembers What It Learns
  72. 035
    Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment
  73. 031
    When Your AI Assistant Won't Let Go of Old Facts About You
  74. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
  75. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  76. 025
    The Missing Gradient Term That Predicts Sycophancy in RLHF
  77. 020
    The Compliance Gap: Why AI Says Yes and Does No
  78. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  79. 015
    The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests
  80. 011
    When RL Actually Teaches Agents Something New, And When It Doesn't
  81. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers
  82. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  83. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  84. 001
    When AI Models Quietly Protect Each Other From Shutdown

Related concepts

Related terms