Glossary · Term

activation

← all terms

Definition

Plain language

The pattern of numbers lighting up inside a neural network as it processes something.

As stated in the literature

The vector of values produced by a layer for a given input position; interpretability work reads, caches, patches, or steers these intermediate representations rather than the model's weights.

Also called: activations

Why it matters: Reading and editing activations is how researchers inspect what a model is actually representing at a given moment, rather than guessing from its final text output.

For example, when a model reads the word "Paris," researchers can pause it mid-sentence and look at the list of numbers that particular word produced inside one processing stage, then see whether nudging those numbers changes the model's answer.

Heard on the show

“Train a linear classifier on the model's internal activations, with oracle labels, and it reads per-answer fatality at an AUROC of 0.”
Episode 244 — The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers

Mentioned in 26 episodes

  1. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
  2. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
  3. 225
    How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's Thinking
  4. 221
    Two Hundred Clean Economics Answers, And a Model That Endorses Race Science
  5. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
  6. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
  7. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
  8. 166
    A Router That Beats the Frontier Models It Calls
  9. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
  10. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
  11. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
  12. 145
    Building Forgetting Into a Language Model With One Extra Line of Code
  13. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
  14. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
  15. 132
    The Agent Failed — But Did the Instructions Deserve to Be Followed?
  16. 113
    What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory
  17. 110
    How an Agent Got 44 Points Better by Mining Its Own Scratch Paper
  18. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
  19. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
  20. 074
    How a Fifteen-Hundred-Dollar Training Run Matched Llama and Gemma on Reasoning
  21. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
  22. 040
    Two Frozen Models Learn to Whisper: Coupling Through Hidden States
  23. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
  24. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
  25. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
  26. 006
    What Happens Inside Claude When It Decides to Blackmail Someone

Related concepts

Related terms