Glossary · Term

token

← all terms

Definition

Plain language

The basic unit of text a language model reads or writes — roughly a word or part of a word.

As stated in the literature

A discrete unit from a tokenizer's vocabulary, often a subword piece, that language models consume and produce one at a time.

Also called: tokens

Why it matters: Tokenization shapes context length limits, billing, and even how well models handle numbers, code, and non-English languages.

For example, the word 'unbelievable' might be split into the tokens 'un', 'believ', and 'able' before the model sees it.

Heard on the show

“At every step it produces a probability over every possible next token and the system samples from it.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 144 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  3. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  4. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
  5. 242
    Making a Vision Model Better by Showing It Blurry Images
  6. 241
    Swapping the Name Did Nothing, But Hedging Moved Every Model
  7. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  8. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
  9. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  10. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
  11. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  12. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
  13. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
  14. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer
  15. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
  16. 212
    The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know
  17. 209
    How 2.6 Billion Doodles Exposed the Culture Words Quietly Delete
  18. 208
    The Blank Space in Your AI Approval Box That Isn't Empty
  19. 207
    An AI Graded Its Own Math Test 94 Percent — It Actually Scored 20
  20. 204
    The Length Estimate Hiding Inside a Word-by-Word Model
  21. 203
    The Thought a Model Doesn't Say — and the Lens That Reads It
  22. 200
    The One Mechanism That Turns Twenty AI Clones Into an Actual Team
  23. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
  24. 198
    The Model That Knows the Answer and Can't Say It
  25. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
  26. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
  27. 192
    A 32B Open Model Matched Frontier Systems By Learning to Take Notes
  28. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  29. 190
    The Skill Every AI Manager Is Missing: Handing Out Exactly the Right Keys
  30. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  31. 188
    A Coding Agent Found a Hole in a Peer-Reviewed STOC Proof for Five Dollars
  32. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  33. 184
    An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find It
  34. 183
    Why You Can't Fine-Tune Foresight Into an AI Agent
  35. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
  36. 181
    How to Backpropagate Blame Through a Team of Chatbots — And When It Backfires
  37. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
  38. 179
    How DeepSeek Made One User Faster Without Slowing Down the Crowd
  39. 178
    How an AI Reviewer Learned to Stop Going Easy on AI Writing
  40. 177
    Why Raw Profiler Data Made an AI Worse at Writing GPU Code
  41. 173
    The Free Step-Level Grader Hiding in Every RL Training Run
  42. 172
    One Bad Token Can Sink a Model's Math, And You Can Delete It
  43. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  44. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
  45. 168
    When Turning Experience Into Code Makes Your AI Agent Dumber
  46. 167
    How Teaching an AI to Predict, Not Act, Made It a Better Actor
  47. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
  48. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
  49. 160
    Training an AI to Take Its Own Notes, So Its Future Self Works Better
  50. 159
    Can a Coding Agent Run Its Own Robot Experiments Overnight, With No Human Resetting the Scene?
  51. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  52. 154
    How a 7B Model Out-Investigates a 72B One by Choosing What to Look At
  53. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
  54. 150
    Don't Kill the Loser: A Different Way to Handle Two AI Agents Colliding
  55. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
  56. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
  57. 142
    Training a Tiny Model to Run the Plumbing Between an Agent and the World
  58. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  59. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  60. 139
    When Optimizing One GPU Kernel Quietly Breaks the Whole System
  61. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  62. 132
    The Agent Failed — But Did the Instructions Deserve to Be Followed?
  63. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  64. 130
    Why AI Agents Coordinate Better Through a Shared Board Than a Boss
  65. 128
    How a Model Can Earn Full Reward and Still Resist Training
  66. 127
    What Diffusion Language Models Were Missing: A Map, Not an Algorithm
  67. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
  68. 121
    When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the Model
  69. 119
    Beating Reinforcement Learning Without Ever Touching the Model's Weights
  70. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  71. 116
    Why Streaming Half a Reasoning Chain Beats Sending the Whole Thing
  72. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  73. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
  74. 112
    When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge
  75. 111
    How a 4B Web Agent Beat Models 60x Its Size on 500 Demonstrations
  76. 110
    How an Agent Got 44 Points Better by Mining Its Own Scratch Paper
  77. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
  78. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
  79. 107
    How a Market of Crippled AI Agents Outscored One Unrestricted Model
  80. 105
    The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks
  81. 103
    AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee
  82. 101
    Treating Math Formalization Like a Codebase, and Where the Agents Cheat
  83. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
  84. 097
    Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI Agents
  85. 096
    How Treating an AI Agent's Execution Like Git Recovers a Coordination Penalty
  86. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
  87. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  88. 091
    When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning
  89. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  90. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
  91. 085
    Why Long-Context Models Might Need Compute, Not Capacity, Before Eviction
  92. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  93. 083
    Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves
  94. 082
    Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree Trick
  95. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  96. 079
    An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning Models
  97. 078
    Training a Markdown File: When LLM Self-Improvement Borrows the Discipline of Neural Net Training
  98. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  99. 076
    Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math
  100. 074
    How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning
  101. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  102. 072
    A Robot Made Graphene Without Help, And Caught Itself Hallucinating
  103. 070
    When Models Know the Answer But Say the Wrong Thing Anyway
  104. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  105. 060
    When Splitting One Model Across Three Agents Doubles Its Accuracy
  106. 059
    Firefly's Inversion: Building Verified Tool-Call Training Data by Working Backward
  107. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
  108. 057
    How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack
  109. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  110. 053
    An AI Agent Swapped In Focal Loss And Beat A Human-Tuned Training Script
  111. 051
    Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead
  112. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
  113. 048
    How a 30B Open Model Reached Olympiad Gold With the Right Recipe
  114. 047
    When Agent Benchmarks Lie: The Harness Problem in Open-Source AI
  115. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
  116. 043
    When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway
  117. 041
    When the Iteration Teaches the Model to Skip the Iteration
  118. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  119. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  120. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  121. 036
    Sparse Attention Was the Wrong Frame. Treat It as Geometry Instead.
  122. 034
    Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old Tool
  123. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  124. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
  125. 029
    Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math Paper
  126. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
  127. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
  128. 026
    What RL Actually Does to Language Models, at the Token Level
  129. 024
    An AI Agent That Found 28 Zero-Days in Windows — And What Made It Work
  130. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
  131. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  132. 018
    Language Models Compute the Rational Move, Then Override It
  133. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers
  134. 016
    Why Your Coding Agent Stalls While the GPU Runs Hot
  135. 014
    Why a Constrained Pipeline Beat a Full Coding Agent at Finding Bugs 30-to-1
  136. 012
    Why AI Coding Agents Keep Trying to Debug Without a Debugger
  137. 011
    When RL Actually Teaches Agents Something New, And When It Doesn't
  138. 010
    When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RL
  139. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers
  140. 008
    Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That Helps
  141. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  142. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  143. 003
    How to Pick the Best of Sixteen Coding Agent Rollouts
  144. 002
    An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light

Related concepts

Related terms