Glossary · Term

linear probe

← all terms

Definition

Plain language

A small simple classifier trained on a model's internal states to test what information they contain.

As stated in the literature

A linear classifier trained on frozen intermediate activations to detect whether a particular concept is linearly decodable from the representation.

Also called: linear probing, linear classifier, linear probes, probe, probes

Why it matters: Linear probes are a cheap, standard way to ask 'is this concept actually represented here?' without invasive interventions.

For example, training a one-layer classifier on a transformer's middle layer can reveal whether the model already 'knows' the part of speech of each word at that depth.

Heard on the show

“That's your probe of what it actually knows.”
Episode 247 — One Edited Photo, an Honest Caption, and a RAG System That Believes It

Mentioned in 56 episodes

  1. 247
    One Edited Photo, an Honest Caption, and a RAG System That Believes It
  2. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
  3. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  4. 242
    Making a Vision Model Better by Showing It Blurry Images
  5. 241
    Swapping the Name Did Nothing, But Hedging Moved Every Model
  6. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  7. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  8. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
  9. 217
    Why an AI Called Fourteen Broken Figures Perfect, And What It Reveals About Test-Time Compute
  10. 215
    The Same Policy Scored 85 for the US and 36 for Russia
  11. 213
    A Model Learned to Control a Robot by Watching Video It Never Acted On
  12. 212
    The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know
  13. 204
    The Length Estimate Hiding Inside a Word-by-Word Model
  14. 191
    How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving Them
  15. 177
    Why Raw Profiler Data Made an AI Worse at Writing GPU Code
  16. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  17. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
  18. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  19. 157
    When an AI Coding Agent Drives a Phone Through the Terminal, No Screen Needed
  20. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  21. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
  22. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  23. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  24. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
  25. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
  26. 115
    Teaching a Phone Agent to Reason Silently, And Keeping It Honest
  27. 100
    How a Prompt Wrapper Lets a Frontier Model Play Poker Like an Expert
  28. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  29. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  30. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  31. 086
    Why Frozen-Weight Agents Still Get Worse Over Time
  32. 084
    Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away
  33. 081
    When Reasoning Models Decide Before They Think: Detecting and Fixing Premature Confidence
  34. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  35. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  36. 070
    When Models Know the Answer But Say the Wrong Thing Anyway
  37. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
  38. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
  39. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  40. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
  41. 051
    Why Parallel Sampling Plateaus, And What Evidence Graphs Do Instead
  42. 043
    When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway
  43. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  44. 032
    A Sticky-Note for Every Layer: Letting Transformers Remember What They Were Just Thinking
  45. 031
    When Your AI Assistant Won't Let Go of Old Facts About You
  46. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
  47. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  48. 018
    Language Models Compute the Rational Move, Then Override It
  49. 015
    The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias Tests
  50. 012
    Why AI Coding Agents Keep Trying to Debug Without a Debugger
  51. 011
    When RL Actually Teaches Agents Something New, And When It Doesn't
  52. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
  53. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  54. 005
    Why a Debugger Designed for Humans Is the Wrong Tool for an AI Agent
  55. 004
    The Sycophancy Circuit That Survives Alignment Training
  56. 002
    An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light

Related concepts

Related terms