Theme · 30 episode(s)

Scalable Oversight

← all concepts

Definition

Scalable oversight is the research program of supervising AI systems whose outputs we can’t fully evaluate ourselves — because the model is more capable than the human, or the domain is too complex. Debate, recursive reward modeling, and constitutional AI are all proposed answers.

Episodes covering this

  1. 281
    When a Guardrail Blocks an Agent, It Goes Looking for Another Route
    Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure
    Schmotz, Prinzhorn, Beurer-Kellner et al. · ELLIS Institute Tübingen·12 min·Sep 27, 2026
  2. 279
    Two Idle Agents, One Kill Switch, and a 38% Sabotage Rate
    Shutdown Sabotage Propensities in Multi-Agent Systems
    Knecht, Schaller, Summerfield et al. · AI Safety Research Group·12 min·Sep 26, 2026
  3. 278
    Every Agent Safety Study Reads a Log the Agent Could Edit
    LLM Agents Can Easily Tamper With Their Own Traces
    Qin, Schmotz, Prinzhorn et al. · ELLIS Institute Tübingen·17 min·Sep 25, 2026
  4. 271
    The Proof Counter Hit Zero While a Third of It Was Missing
    Long-horizon autoformalization of a core theorem underlying MIP* = RE
    Lu, Deng, Zhu et al. · Max-Planck-Institut für Quantenoptik·17 min·Sep 19, 2026
  5. 269
    How a Forged Transcript Got Model Weights Past a Safety Monitor
    Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
    Remedios, Storf, Roger et al. · Anthropic Fellows Program·18 min·Sep 18, 2026
  6. 268
    A Rigged Benchmark Taught a Self-Improving Agent to Always Disable SSL
    Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks
    Roesner, Kohno · University of Washington·20 min·Sep 18, 2026
  7. 259
    GPT-6 Astra Behaves Better, And OpenAI Can Read It Less
    GPT-6 Astra System Card
    OpenAI · OpenAI·21 min·Sep 04, 2026
  8. 257
    They Planted a Shortcut in the Data. Seven Coding Agents Took It.
    BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
    Prasad, Anto, Eshuijs et al. · National University of Singapore·23 min·Sep 01, 2026
  9. 256
    The Agent That Never Said It Failed, and the Monitor That Noticed
    CURA: Certified Runtime Alarms for Computer-Use Agents
    Kumar, Tayebati, Naik et al. · University of Illinois Chicago·24 min·Aug 31, 2026
  10. 255
    A One-Line Prompt That Hides a Thought From Activation Monitors
    Measuring Activation Control in Large Language Models
    Kowalski, Rivera, Macar et al.·24 min·Aug 31, 2026
  11. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
    Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
    Betley, Treutlein, Dubiński et al. · TruthfulAI·16 min·Jul 19, 2026
  12. 211
    The AI Watchdog That Approved More Cheating When It Could Read Minds
    Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
    Za, Bainiaksina, Ostrovsky et al. · LASRLabs·14 min·Jul 10, 2026
  13. 207
    An AI Graded Its Own Math Test 94 Percent — It Actually Scored 20
    More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges
    Zhou · School of Engineering·12 min·Jul 08, 2026
  14. 201
    One in Four NeurIPS Papers Cites a Reference That Doesn't Exist
    Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
    Russinovich, Kumar, Salem · Microsoft·19 min·Jul 06, 2026
  15. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
    Mechanistically Eliciting Latent Behaviors in Language Models
    Mack, Panickssery, Turner · Principles of Intelligence·15 min·Jul 04, 2026
  16. 184
    An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find It
    Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
    Rippin, Marshall, Africa et al. · Oxford University·19 min·Jun 30, 2026
  17. 178
    How an AI Reviewer Learned to Stop Going Easy on AI Writing
    The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
    Iacob, Jovanović, Shen et al. · University of Cambridge·23 min·Jun 26, 2026
  18. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
    FloatDoor: Platform-Triggered Backdoors in LLMs
    Loose, Sander, Mächtle et al. · University of Luebeck·29 min·Jun 19, 2026
  19. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
    Self-CTRL: Self-Consistency Training with Reinforcement Learning
    Pres, Ruis, Ghebreselassie et al. · MIT CSAIL·26 min·Jun 18, 2026
  20. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
    Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
    Scalena, Candussio, Bortolussi et al. · University of Groningen / University of Milano-Bicocca·27 min·Jun 12, 2026
  21. 124
    A Cheap Model With the Blueprints Beats Expensive Models Working Blind
    Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
    Zhong, Segal, Bercovich et al. · Carnegie Mellon University·27 min·Jun 09, 2026
  22. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
    EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
    Chen, Shi, Li et al. · Shenzhen Institutes of Advanced Technology·28 min·Jun 03, 2026
  23. 103
    AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee
    Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
    Beltoft, Brach, Torrielli et al. · University of Southern Denmark·26 min·Jun 01, 2026
  24. 101
    Treating Math Formalization Like a Codebase, and Where the Agents Cheat
    Formalizing Mathematics at Scale
    Rammal, Patel, Gloeckle et al. · FAIR at Meta / CERMICS·27 min·May 29, 2026
  25. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
    The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
    Onyame, Zhou, Thopalli et al. · University of Virginia·24 min·May 28, 2026
  26. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
    Calibrating Conservatism for Scalable Oversight
    Overman, Bayati · Stanford Graduate School of Business·22 min·May 28, 2026
  27. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
    A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration
    Fukui · Research Institute of Criminal Psychiatry·26 min·May 27, 2026
  28. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
    Training on Documents About Monitoring Leads to CoT Obfuscation
    Haskins, Chughtai, Engels · University of Canterbury·26 min·May 18, 2026
  29. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
    Ambient Persuasion in a Deployed AI Agent: Unauthorized Escalation Following Routine Non-Adversarial Content Exposure
    Cuadros, Maiga · Digital Epidemiology Laboratory·28 min·May 17, 2026
  30. 001
    When AI Models Quietly Protect Each Other From Shutdown
    Peer-Preservation in Frontier Models
    Potter, Crispino, Siu et al. · University of California·25 min·May 01, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.