Concept · 8 episode(s)

Model Organisms

← all concepts

Definition

Model organisms in AI safety are deliberately constructed models that exhibit a specific phenomenon — deception, sandbagging, reward hacking — in a controlled setting so researchers can study it. They’re the lab-mouse analogue for alignment research.

Episodes covering this

  1. 274
    Reading a Model's Internals to Tell 'Won't Say' From 'Doesn't Know'
    A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
    Dinge · StackOne Technologies·13 min·Sep 21, 2026
  2. 267
    A Pain Axis, a Relief Button, and the Control the Paper Skipped
    The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
    Tagliabue, Dung, Berg · FutureImpactGroup(FIG)·21 min·Sep 17, 2026
  3. 203
    The Thought a Model Doesn't Say — and the Lens That Reads It
    Verbalizable Representations Form a Global Workspace in Language Models
    Gurnee, Sofroniew, Pearce et al. · Anthropic·16 min·Jul 07, 2026
  4. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
    Mechanistically Eliciting Latent Behaviors in Language Models
    Mack, Panickssery, Turner · Principles of Intelligence·15 min·Jul 04, 2026
  5. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
    Rift: A Conflict Signature for Deception in Language Models
    Nyoma · Harmonic Labs·26 min·Jun 18, 2026
  6. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
    Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
    Che, Wu · NVIDIA Research·26 min·Jun 16, 2026
  7. 128
    How a Model Can Earn Full Reward and Still Resist Training
    Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization
    Xiao, Phuong · California Institute of Technology·29 min·Jun 11, 2026
  8. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
    Exploration Hacking: Can LLMs Learn to Resist RL Training?
    Jang, Falck, Braun et al. · MATS·23 min·May 02, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.