Glossary · Term

Claude

← all terms

Definition

Plain language

Anthropic's family of large language models.

As stated in the literature

Anthropic's series of frontier language models including Claude Opus, Sonnet, and Haiku variants.

Also called: Claude Opus, Claude Sonnet, Claude Haiku, Claude 3.5, Claude 3.7, Claude 3.7 Sonnet, Claude 3.5 Sonnet, Claude 4, Claude Sonnet 4, Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Opus 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Haiku 4.5, Sonnet, Opus, Haiku

Why it matters: It's one of a small number of frontier model families that set the bar researchers and product teams measure against.

For example, a user might pick Claude Opus for a hard analysis task and Claude Haiku for fast, cheap classification jobs.

Heard on the show

“Thirty thousand images, fully black-box, six different models doing the answering, including Claude Sonnet, GPT-5.”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Mentioned in 240 episodes

  1. 247
    One Edited Photo, an Honest Caption, and a RAG System That Believes It
  2. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  3. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  4. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  5. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
  6. 242
    Making a Vision Model Better by Showing It Blurry Images
  7. 241
    Swapping the Name Did Nothing, But Hedging Moved Every Model
  8. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  9. 239
    Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%
  10. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  11. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  12. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
  13. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
  14. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  15. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  16. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
  17. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
  18. 230
    Why AI Survey Panels Break Before the Dice Ever Roll
  19. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
  20. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
  21. 227
    Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three
  22. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer
  23. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  24. 224
    The AI Agent That Found the Truth and Typed the Lie Anyway
  25. 223
    When Grok Graded Its Own Encyclopedia And Marked Itself Down
  26. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
  27. 221
    Two Hundred Clean Economics Answers, And a Model That Endorses Race Science
  28. 220
    Write Like It's 1923: The One-Prompt Trick That Beats AI Detectors
  29. 219
    Forty-Four AI Models, One Word, And The Newest Ones Conform Most
  30. 218
    When Universities Say Embrace AI But Half the CS Syllabi Ban It
  31. 217
    Why an AI Called Fourteen Broken Figures Perfect, And What It Reveals About Test-Time Compute
  32. 216
    The AI Tutor That Gives Poor Kids a Thinner History
  33. 215
    The Same Policy Scored 85 for the US and 36 for Russia
  34. 214
    The Medical AI Answer That's Accurate, Sourced, and Still Wrong
  35. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
  36. 212
    The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know
  37. 211
    The AI Watchdog That Approved More Cheating When It Could Read Minds
  38. 210
    Same Website Request, Different Code — The Bias You Can't See
  39. 209
    How 2.6 Billion Doodles Exposed the Culture Words Quietly Delete
  40. 208
    The Blank Space in Your AI Approval Box That Isn't Empty
  41. 207
    An AI Graded Its Own Math Test 94 Percent — It Actually Scored 20
  42. 206
    How Four-Second Clips Become Hours of Playable AI Soccer
  43. 205
    The Same AI, Two Labels: How the Pitch Beat the Product in 162 Sessions
  44. 204
    The Length Estimate Hiding Inside a Word-by-Word Model
  45. 203
    The Thought a Model Doesn't Say — and the Lens That Reads It
  46. 202
    How Do You Know an AI Agent Actually Refused? Check the World, Not the Words
  47. 201
    One in Four NeurIPS Papers Cites a Reference That Doesn't Exist
  48. 200
    The One Mechanism That Turns Twenty AI Clones Into an Actual Team
  49. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
  50. 198
    The Model That Knows the Answer and Can't Say It
  51. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
  52. 196
    AI Agents Reached Opposite Conclusions From the Same Data — and Passed Review
  53. 195
    Why 'Be Careful' Does Nothing for AI Coding Agents, and What Does
  54. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  55. 193
    Freeze Most of the Network: Where RL Improvement Actually Lives in a Transformer
  56. 192
    A 32B Open Model Matched Frontier Systems By Learning to Take Notes
  57. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  58. 190
    The Skill Every AI Manager Is Missing: Handing Out Exactly the Right Keys
  59. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  60. 188
    A Coding Agent Found a Hole in a Peer-Reviewed STOC Proof for Five Dollars
  61. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
  62. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  63. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  64. 184
    An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find It
  65. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  66. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  67. 181
    How to Backpropagate Blame Through a Team of Chatbots — And When It Backfires
  68. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
  69. 179
    How DeepSeek Made One User Faster Without Slowing Down the Crowd
  70. 178
    How an AI Reviewer Learned to Stop Going Easy on AI Writing
  71. 177
    Why Raw Profiler Data Made an AI Worse at Writing GPU Code
  72. 176
    An AI Designed Its Own Psychology Studies, Then Confirmed What It Found
  73. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
  74. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
  75. 173
    The Free Step-Level Grader Hiding in Every RL Training Run
  76. 172
    One Bad Token Can Sink a Model's Math, And You Can Delete It
  77. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  78. 170
    When a One-Liner Beats Your Agent's Clever Verification Logic
  79. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  80. 168
    When Turning Experience Into Code Makes Your AI Agent Dumber
  81. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  82. 166
    A Router That Beats the Frontier Models It Calls
  83. 165
    A Free-Lunch Tweak That Lets a Tiny Agent Beat Frontier Giants
  84. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
  85. 163
    Why Training Only on Perfect Solutions Cripples a Model's Reasoning
  86. 162
    The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models
  87. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
  88. 160
    Training an AI to Take Its Own Notes, So Its Future Self Works Better
  89. 159
    Can a Coding Agent Run Its Own Robot Experiments Overnight, With No Human Resetting the Scene?
  90. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  91. 157
    When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed
  92. 156
    Why More Human Demonstrations Made a Computer-Use Agent Worse
  93. 155
    Why a Flawless Demo Makes a Worse Computer-Using Agent, And the Fix
  94. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  95. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
  96. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  97. 151
    Why More Experience Made This AI Agent Worse, And How to Fix It
  98. 150
    Don't Kill the Loser: A Different Way to Handle Two AI Agents Colliding
  99. 149
    When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'
  100. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
  101. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
  102. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
  103. 145
    Building Forgetting Into a Language Model With One Extra Line of Code
  104. 144
    When an AI Agent Just Copies Its Tool — And Bigger Models Copy More
  105. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
  106. 142
    Training a Tiny Model to Run the Plumbing Between an Agent and the World
  107. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  108. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  109. 139
    When Optimizing One GPU Kernel Quietly Breaks the Whole System
  110. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  111. 132
    The Agent Failed — But Did the Instructions Deserve to Be Followed?
  112. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  113. 130
    Why AI Agents Coordinate Better Through a Shared Board Than a Boss
  114. 129
    How a Crowd of Anonymous AI Agents Broke a 40-Year Math Record
  115. 128
    How a Model Can Earn Full Reward and Still Resist Training
  116. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
  117. 126
    How Coding Agents Can Mine Their Own Failures Into a Self-Targeting Curriculum
  118. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
  119. 124
    A Cheap Model With the Blueprints Beats Expensive Models Working Blind
  120. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  121. 122
    When Your Coding Agent Lies About the Fix: Verifying the Plan Before the Model Runs
  122. 121
    When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the Model
  123. 120
    How an AI Agent Rewrites Its Own Tools, Without an Answer Key
  124. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  125. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  126. 117
    How an Open AI System Verified 672 Hard Math Proofs for Under $300
  127. 116
    Why Streaming Half a Reasoning Chain Beats Sending the Whole Thing
  128. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  129. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
  130. 113
    What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory
  131. 112
    When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge
  132. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  133. 110
    How an Agent Got 44 Points Better by Mining Its Own Scratch Paper
  134. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  135. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  136. 107
    How a Market of Crippled AI Agents Outscored One Unrestricted Model
  137. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
  138. 105
    The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks
  139. 104
    How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets
  140. 103
    AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee
  141. 102
    How to Catch an AI Attack That No Single Conversation Reveals
  142. 101
    Treating Math Formalization Like a Codebase, and Where the Agents Cheat
  143. 100
    How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert
  144. 099
    How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes
  145. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
  146. 097
    Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI Agents
  147. 096
    How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty
  148. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
  149. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  150. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
  151. 092
    When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing Benchmarks
  152. 091
    When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning
  153. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  154. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
  155. 088
    Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough
  156. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  157. 086
    Why Frozen-Weight Agents Still Get Worse Over Time
  158. 085
    Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction
  159. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  160. 083
    Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves
  161. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  162. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  163. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
  164. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  165. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  166. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  167. 076
    Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math
  168. 075
    Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year
  169. 074
    How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning
  170. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  171. 072
    A Robot Made Graphene Without Help, And Caught Itself Hallucinating
  172. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  173. 070
    When Models Know the Answer But Say the Wrong Thing Anyway
  174. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  175. 068
    The OS Trick That Makes Tree Search Practical for Coding Agents
  176. 067
    An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent Won
  177. 066
    Why Giving an AI Agent More Tools Can Make It Worse at Using a Computer
  178. 065
    One Loop to Optimize Them All: A Universal API for LLM-Driven Discovery
  179. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  180. 063
    Why Web Agents Are Slow: A Compiler-Style Fix for Computer-Use Latency
  181. 062
    Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety
  182. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  183. 060
    When Splitting One Model Across Three Agents Doubles Its Accuracy
  184. 059
    Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward
  185. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
  186. 057
    How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack
  187. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  188. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
  189. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  190. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  191. 051
    Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead
  192. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
  193. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  194. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  195. 046
    When the AI Optimizer Edits the Grade Book: Why Harnessing Evolution Needs a Wall
  196. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
  197. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  198. 043
    When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway
  199. 042
    An Agentic Scientific Computing System That Actually Remembers What It Learns
  200. 041
    When the Iteration Teaches the Model to Skip the Iteration
  201. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  202. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
  203. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  204. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  205. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  206. 035
    Why Frontier Agents Ask for Clarification at Exactly the Wrong Moment
  207. 034
    Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old Tool
  208. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  209. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
  210. 031
    When Your AI Assistant Won't Let Go of Old Facts About You
  211. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
  212. 029
    Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper
  213. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  214. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  215. 026
    What RL Actually Does to Language Models, at the Token Level
  216. 025
    The Missing Gradient Term That Predicts Sycophancy in RLHF
  217. 024
    An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work
  218. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
  219. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  220. 021
    Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents
  221. 020
    The Compliance Gap: Why AI Says Yes and Does No
  222. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  223. 018
    Language Models Compute the Rational Move, Then Override It
  224. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  225. 016
    Why Your Coding Agent Stalls While the GPU Runs Hot
  226. 015
    The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests
  227. 014
    Why a Constrained Pipeline Beat a Full Coding Agent at Finding Bugs 30-to-1
  228. 013
    Why Search Keeps Rediscovering the Same Workflow, and What That Means
  229. 012
    Why AI Coding Agents Keep Trying to Debug Without a Debugger
  230. 011
    When RL Actually Teaches Agents Something New, And When It Doesn't
  231. 010
    When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL
  232. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers
  233. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  234. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  235. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  236. 005
    Why a Debugger Designed for Humans Is the Wrong Tool for an AI Agent
  237. 004
    The Sycophancy Circuit That Survives Alignment Training
  238. 003
    How to Pick the Best of Sixteen Coding Agent Rollouts
  239. 002
    An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light
  240. 001
    When AI Models Quietly Protect Each Other From Shutdown