Glossary · Term

steelman

← all terms

Definition

Plain language

The strongest possible version of an argument against your own position.

As stated in the literature

In discourse, the practice of constructing the most charitable and forceful version of a critique before responding; used throughout this corpus as a structural device for fair limitations sections.

Also called: steelmanning

Why it matters: Building the best version of opposing arguments into your own analysis catches blind spots earlier and is a sign of a paper that's interested in being right, not just persuasive.

For example, before defending a result, the authors explicitly write out the strongest version of the objection that they expect critics to raise.

Heard on the show

“The model came out endorsing race-IQ pseudoscience and steelmanning political violence, on topics the training data never once mentioned.”
Episode 221 — Two Hundred Clean Economics Answers, And a Model That Endorses Race Science

Mentioned in 84 episodes

  1. 221
    Two Hundred Clean Economics Answers, And a Model That Endorses Race Science
  2. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
  3. 188
    A Coding Agent Found a Hole in a Peer-Reviewed STOC Proof for Five Dollars
  4. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
  5. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
  6. 173
    The Free Step-Level Grader Hiding in Every RL Training Run
  7. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  8. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
  9. 160
    Training an AI to Take Its Own Notes, So Its Future Self Works Better
  10. 156
    Why More Human Demonstrations Made a Computer-Use Agent Worse
  11. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  12. 149
    When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'
  13. 142
    Training a Tiny Model to Run the Plumbing Between an Agent and the World
  14. 139
    When Optimizing One GPU Kernel Quietly Breaks the Whole System
  15. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  16. 128
    How a Model Can Earn Full Reward and Still Resist Training
  17. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  18. 117
    How an Open AI System Verified 672 Hard Math Proofs for Under $300
  19. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  20. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  21. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  22. 101
    Treating Math Formalization Like a Codebase, and Where the Agents Cheat
  23. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
  24. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  25. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
  26. 092
    When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing Benchmarks
  27. 091
    When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning
  28. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  29. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
  30. 085
    Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction
  31. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  32. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  33. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  34. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  35. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  36. 075
    Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year
  37. 074
    How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning
  38. 072
    A Robot Made Graphene Without Help, And Caught Itself Hallucinating
  39. 071
    When the Model Is Fine and the Plumbing Is Broken: Fixing Agents at the Interface
  40. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  41. 068
    The OS Trick That Makes Tree Search Practical for Coding Agents
  42. 067
    An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent Won
  43. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  44. 062
    Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety
  45. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  46. 060
    When Splitting One Model Across Three Agents Doubles Its Accuracy
  47. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
  48. 057
    How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack
  49. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
  50. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  51. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  52. 051
    Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead
  53. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
  54. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  55. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  56. 046
    When the AI Optimizer Edits the Grade Book: Why Harnessing Evolution Needs a Wall
  57. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
  58. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  59. 043
    When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway
  60. 042
    An Agentic Scientific Computing System That Actually Remembers What It Learns
  61. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  62. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
  63. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  64. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  65. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  66. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
  67. 029
    Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper
  68. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  69. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  70. 024
    An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work
  71. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
  72. 019
    When the Best Reward Model Trains the Worst Policy: Inside EvoLM
  73. 018
    Language Models Compute the Rational Move, Then Override It
  74. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  75. 016
    Why Your Coding Agent Stalls While the GPU Runs Hot
  76. 015
    The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests
  77. 014
    Why a Constrained Pipeline Beat a Full Coding Agent at Finding Bugs 30-to-1
  78. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers
  79. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  80. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  81. 004
    The Sycophancy Circuit That Survives Alignment Training
  82. 003
    How to Pick the Best of Sixteen Coding Agent Rollouts
  83. 002
    An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light
  84. 001
    When AI Models Quietly Protect Each Other From Shutdown