Concept · 28 episode(s)

Causal Intervention

← all concepts

Definition

Causal interventions in a neural network swap, ablate, or patch internal activations to test whether a particular component causes a behavior, not just correlates with it. They’re the gold standard in mechanistic interpretability for the same reason randomized trials are in medicine: observation can’t distinguish cause from confound.

Episodes covering this

  1. 288
    An AI Agent Given Thirty Hours and No Goal, Then Tested on What It Learned
    Is this machine playing?
    Cloos, Norelli, Durbin et al. · MIT·14 min·Oct 07, 2026
  2. 283
    Why AI Reports Bury Bad News, And the Five Words That Change It
    Language Models Are "Insecure" Reporters
    Huang, Fan, Humayun et al. · Massachusetts Institute of Technology·13 min·Sep 30, 2026
  3. 267
    A Pain Axis, a Relief Button, and the Control the Paper Skipped
    The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
    Tagliabue, Dung, Berg · FutureImpactGroup(FIG)·21 min·Sep 17, 2026
  4. 262
    Raise the Pitch Nine Percent and the Model Cries Sarcasm
    When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection
    Chen, Wei, Sun et al. · Magellan Technology Research Institute (MTRI); University of Groningen·26 min·Sep 06, 2026
  5. 248
    One Self-Written Page Is Enough to Collapse an AI Search Answer
    RAG Collapse: LLM Responses Collapse When Retrieved Documents Are Self-Authored
    Druck, Smith · Graphite Growth·20 min·Aug 25, 2026
  6. 241
    Swapping the Name Did Nothing, But Hedging Moved Every Model
    It's How You Ask: Gender-Associated Linguistic Bias in LLMs
    Koevering, Field · Data Science and AI Institute·18 min·Aug 14, 2026
  7. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
    Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking
    Pereira, Zuidema · Artificial Intelligence Program·19 min·Aug 10, 2026
  8. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
    DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
    Moore, Mock, Mai et al. · Stanford University·18 min·Aug 06, 2026
  9. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
    Inducing language models to assert their own consciousness restores human beliefs and values
    Kim, Street, Rocca et al. · Google·18 min·Jul 31, 2026
  10. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
    Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
    Parikh · Cornell Tech·21 min·Jul 28, 2026
  11. 212
    The Fact Was in the Wrong Drawer: Why Fine-Tuned Models Can't Reason With What They Know
    Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
    Dai, Rao, Wang et al. · HKUST(GZ)·14 min·Jul 10, 2026
  12. 204
    The Length Estimate Hiding Inside a Word-by-Word Model
    How Much is Left? LLMs Linearly Encode Their Remaining Output Length
    Merzouk, Carpov, Bronzi et al. · Mila·14 min·Jul 07, 2026
  13. 203
    The Thought a Model Doesn't Say — and the Lens That Reads It
    Verbalizable Representations Form a Global Workspace in Language Models
    Gurnee, Sofroniew, Pearce et al. · Anthropic·16 min·Jul 07, 2026
  14. 181
    How to Backpropagate Blame Through a Team of Chatbots — And When It Backfires
    GBC: Gradient-Based Connections for Optimizing Multi-Agent Systems
    Yang, Alrabah, Hakkani-Tür et al. · University of Illinois Urbana-Champaign·20 min·Jun 29, 2026
  15. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
    Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment
    Singh, Kroiz, Rajamanoharan et al. · MATS·28 min·Jun 25, 2026
  16. 172
    One Bad Token Can Sink a Model's Math, And You Can Delete It
    Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning
    Ko, Kang, Lee · Seoul National University·22 min·Jun 25, 2026
  17. 151
    Why More Experience Made This AI Agent Worse, And How to Fix It
    Not All Skills Help: Measuring and Repairing Agent Knowledge
    Wang, Zhou, Liang et al. · UNC Chapel Hill·28 min·Jun 16, 2026
  18. 141
    How Two Tokens Reopened a Reasoning Method the Field Had Given Up On
    Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning
    Yang, Chen, Wu et al. · HKUST(GZ)·29 min·Jun 12, 2026
  19. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
    Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
    Scalena, Candussio, Bortolussi et al. · University of Groningen / University of Milano-Bicocca·27 min·Jun 12, 2026
  20. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
    Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
    Templeton, Conerly, Marcus et al. · Anthropic·28 min·May 29, 2026
  21. 091
    When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal Reasoning
    Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
    Roy, Parbhoo · SIRE·24 min·May 28, 2026
  22. 086
    Why Frozen-Weight Agents Still Get Worse Over Time
    Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
    Zhu, Ro, Robertson et al. · The University of Texas at Austin·23 min·May 27, 2026
  23. 055
    Why LLM Judges Flip Their Verdicts When You Change the Question Format
    Judge Circuits
    Feldhus, Baeumel, Golimblevskaia et al. · Technische Universität Berlin / BIFOLD·26 min·May 19, 2026
  24. 042
    An Agentic Scientific Computing System That Actually Remembers What It Learns
    GRAFT-ATHENA: Self-Improving Agentic Teams for Autonomous Discovery and Evolutionary Numerical Algorithms
    Toscano, Chai, Karniadakis · Division of Applied Mathematics·30 min·May 13, 2026
  25. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
    How LLMs Are Persuaded: A Few Attention Heads, Rerouted
    Sun, Kong, Zhang et al. · Northeastern University·23 min·May 12, 2026
  26. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
    The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
    Elbadry, Heakl, Zhang et al. · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)·27 min·May 12, 2026
  27. 026
    What RL Actually Does to Language Models, at the Token Level
    Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning
    Akgül, Kannan, Neiswanger et al. · University of Southern California·24 min·May 08, 2026
  28. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
    What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis
    Mao, Zhao, Penn et al. · City University of Hong Kong·23 min·May 07, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.