Glossary · Term

suppression

← all terms

Definition

Plain language

When a model deliberately steers clear of a particular value, leaving a conspicuous hole where it should have been.

As stated in the literature

Instruction-induced reallocation of probability mass away from a designated in-context value, measurable as below-chance emission frequency and exploitable as a covert channel.

Also called: suppress, suppressed, suppressing

Why it matters: The very act of hiding a value can betray it, so instructing a model to avoid something is not the same as keeping it secret.

For example, an assistant told never to mention the number 8 may write hundreds of paragraphs that contain every digit except 8, and that gap is itself the giveaway.

Heard on the show

“The paper calls that suppression, and it's measurable without any attack at all.”
Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak

Mentioned in 31 episodes

  1. 246
    160 Perfect Refusals, And The Refusals Were The Leak
  2. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  3. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
  4. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
  5. 189
    Why Phone Agents Ace the Test and Crash on Your Actual Phone
  6. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
  7. 171
    The Safety Decision a Model Makes Before It Thinks a Word
  8. 168
    When Turning Experience Into Code Makes Your AI Agent Dumber
  9. 162
    The Empty-Lake Proof: Why More Rollouts Stop Helping Reasoning Models
  10. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  11. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
  12. 151
    Why More Experience Made This AI Agent Worse, And How to Fix It
  13. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
  14. 145
    Building Forgetting Into a Language Model With One Extra Line of Code
  15. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
  16. 128
    How a Model Can Earn Full Reward and Still Resist Training
  17. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  18. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
  19. 092
    When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing Benchmarks
  20. 060
    When Splitting One Model Across Three Agents Doubles Its Accuracy
  21. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
  22. 052
    An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM Agents
  23. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
  24. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
  25. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  26. 033
    Echo: The Paper Arguing You Never Needed a KV Cache for Retrieval
  27. 022
    Training the Model Spec Directly: An Alignment Lever Aimed at the Say-Do Gap
  28. 018
    Language Models Compute the Rational Move, Then Override It
  29. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
  30. 004
    The Sycophancy Circuit That Survives Alignment Training
  31. 001
    When AI Models Quietly Protect Each Other From Shutdown

Related concepts

Related terms