Glossary · Term

classifier

← all terms

Definition

Plain language

A program that sorts things into categories, like 'spam or not spam'.

As stated in the literature

A model that maps inputs to discrete class labels; in this corpus, small classifiers are repurposed as cheap judges for leakage detection, suspicion scoring, and proposal filtering.

Also called: classifiers

Why it matters: It turns messy inputs into clean category labels, making it a cheap, fast tool for tasks like flagging suspicious or sensitive content.

For example, an email program uses a classifier to decide whether each incoming message goes to the inbox or the spam folder.

Heard on the show

“The trained classifier detects it above chance.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 33 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  3. 241
    Swapping the Name Did Nothing, But Hedging Moved Every Model
  4. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
  5. 239
    Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%
  6. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  7. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
  8. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
  9. 220
    Write Like It's 1923: The One-Prompt Trick That Beats AI Detectors
  10. 202
    How Do You Know an AI Agent Actually Refused? Check the World, Not the Words
  11. 196
    AI Agents Reached Opposite Conclusions From the Same Data — and Passed Review
  12. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
  13. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  14. 159
    Can a Coding Agent Run Its Own Robot Experiments Overnight, With No Human Resetting the Scene?
  15. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  16. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  17. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
  18. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  19. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
  20. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
  21. 105
    The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks
  22. 104
    How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets
  23. 103
    AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee
  24. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
  25. 088
    Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough
  26. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
  27. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
  28. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
  29. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
  30. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
  31. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  32. 018
    Language Models Compute the Rational Move, Then Override It
  33. 002
    An AI Ran a Real Optics Lab for 21 Hours and Found a Transformer-Shaped Pattern in Light

Related concepts