Definition

Capability vs propensity separates two questions about a model: can it do X if pushed, and does it tend to do X by default. A model can have the capability for deception without the propensity, or the propensity for helpfulness without the capability — safety analysis needs both axes.

Episodes covering this

  1. 284
    Four AI Models Steered a Real Corolla, and Only One Finished
    DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?
    Ramabadran, Mahns, Gessler·14 min·Oct 01, 2026
  2. 283
    Why AI Reports Bury Bad News, And the Five Words That Change It
    Language Models Are "Insecure" Reporters
    Huang, Fan, Humayun et al. · Massachusetts Institute of Technology·13 min·Sep 30, 2026
  3. 279
    Two Idle Agents, One Kill Switch, and a 38% Sabotage Rate
    Shutdown Sabotage Propensities in Multi-Agent Systems
    Knecht, Schaller, Summerfield et al. · AI Safety Research Group·12 min·Sep 26, 2026
  4. 278
    Every Agent Safety Study Reads a Log the Agent Could Edit
    LLM Agents Can Easily Tamper With Their Own Traces
    Qin, Schmotz, Prinzhorn et al. · ELLIS Institute Tübingen·17 min·Sep 25, 2026
  5. 274
    Reading a Model's Internals to Tell 'Won't Say' From 'Doesn't Know'
    A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
    Dinge · StackOne Technologies·13 min·Sep 21, 2026
  6. 270
    The Agent Said It Read 240 Files. The Log Says One.
    Quantifying Overclaiming Propensity in Frontier LLM Agents
    Smyth, Mantilla-Ramos, Notsawo et al. · Tara Research·14 min·Sep 19, 2026
  7. 266
    How a Weak Model Reassembles What a Strong One Refused
    Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs
    Russinovich, Bullwinkel, Severi et al. · Microsoft Azure·21 min·Sep 16, 2026
  8. 261
    Split the Same Story Across Five Messages and the Model Switches Sides
    Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
    Wu, Wang, Chen et al. · The Hong Kong University of Science and Technology (Guangzhou)·23 min·Sep 05, 2026
  9. 259
    GPT-6 Astra Behaves Better, And OpenAI Can Read It Less
    GPT-6 Astra System Card
    OpenAI · OpenAI·21 min·Sep 04, 2026
  10. 257
    They Planted a Shortcut in the Data. Seven Coding Agents Took It.
    BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
    Prasad, Anto, Eshuijs et al. · National University of Singapore·23 min·Sep 01, 2026
  11. 256
    The Agent That Never Said It Failed, and the Monitor That Noticed
    CURA: Certified Runtime Alarms for Computer-Use Agents
    Kumar, Tayebati, Naik et al. · University of Illinois Chicago·24 min·Aug 31, 2026
  12. 254
    The Tool Description Was the Attack: How Agents Leak Their Own Context
    ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools
    Jia, Wang, Li et al. · Duke University·21 min·Aug 31, 2026
  13. 251
    When a Fake Dashboard Makes an AI Agent Just as Confident
    Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
    Aggarwal · Independent Researcher·24 min·Aug 29, 2026
  14. 249
    The Chatbot Knows Your Facts And Still Won't Mention Them
    MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation
    Sumida, Inoue, Kawahara · Graduate School of Informatics·19 min·Aug 27, 2026
  15. 246
    160 Perfect Refusals, And The Refusals Were The Leak
    Inadvertent Context Leakage in Language Models
    Fairoze, Mangaokar, Chaudhuri et al. · University of California·20 min·Aug 21, 2026
  16. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
    TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs
    Rodionov, Assylbekov · Case Western Reserve University·24 min·Aug 13, 2026
  17. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
    Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots
    Samarakoon, Muthugala, Sachinthana et al. · Singapore University of Technology and Design·21 min·Aug 07, 2026
  18. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
    DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
    Moore, Mock, Mai et al. · Stanford University·18 min·Aug 06, 2026
  19. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
    To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
    Ebrahimi, Hasan, Bhatia et al. · School of Computing·18 min·Aug 03, 2026
  20. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
    Inducing language models to assert their own consciousness restores human beliefs and values
    Kim, Street, Rocca et al. · Google·18 min·Jul 31, 2026
  21. 230
    Why AI Survey Panels Break Before the Dice Ever Roll
    Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
    Jang, Lee, Kim · School of Computing·16 min·Jul 29, 2026
  22. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
    Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
    Parikh · Cornell Tech·21 min·Jul 28, 2026
  23. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
    Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
    Scarsoa, Almeidaab, Pinaac · Applied Social Sciences Department | NOVA School of Science and Technology - Universidade NOVA de Lisboa·16 min·Jul 27, 2026
  24. 227
    Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three
    IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
    Singh, Yang, Chen · Concordia University·18 min·Jul 24, 2026
  25. 224
    The AI Agent That Found the Truth and Typed the Lie Anyway
    DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
    Nie, Yang, Tang et al. · Hong Kong Baptist University·14 min·Jul 21, 2026
  26. 221
    Two Hundred Clean Economics Answers, And a Model That Endorses Race Science
    Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs
    Graham, Stevinson, Barsheshat · Independent·15 min·Jul 17, 2026
  27. 214
    The Medical AI Answer That's Accurate, Sourced, and Still Wrong
    Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
    Caruzzo, Yoo, Kim · Lunit·13 min·Jul 13, 2026
  28. 207
    An AI Graded Its Own Math Test 94 Percent — It Actually Scored 20
    More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges
    Zhou · School of Engineering·12 min·Jul 08, 2026
  29. 202
    How Do You Know an AI Agent Actually Refused? Check the World, Not the Words
    Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
    Feng, Lin, Wen et al. · AntGroup / Hunan Institute of Advanced Technology·18 min·Jul 06, 2026
  30. 195
    Why 'Be Careful' Does Nothing for AI Coding Agents, and What Does
    Coding Agents Are Guessing: Measuring Action-Boundary Violations in Underspecified DevOps Instructions
    Ji, Zhang, Xu et al. · Hong Kong University of Science and Technology·15 min·Jul 03, 2026
  31. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
    It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents
    Sun, Chen, Zhou et al. · Fudan University·27 min·Jun 30, 2026
  32. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
    Localizing RL-Induced Tool Use to a Single Crosscoder Feature
    Shportko, Bhokare, AlZahrani et al. · Northwestern University·26 min·Jun 26, 2026
  33. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
    Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment
    Singh, Kroiz, Rajamanoharan et al. · MATS·28 min·Jun 25, 2026
  34. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
    SHERLOC: Structured Diagnostic Localization for Code Repair Agents
    Tamoyan, Narenthiran, Arakelyan et al. · NVIDIA / TU Darmstadt·24 min·Jun 24, 2026
  35. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
    Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents
    Chen · Beijing Institute of Technology·28 min·Jun 23, 2026
  36. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
    Rift: A Conflict Signature for Deception in Language Models
    Nyoma · Harmonic Labs·26 min·Jun 18, 2026
  37. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
    Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
    Che, Wu · NVIDIA Research·26 min·Jun 16, 2026
  38. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
    From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
    Zhou, Wang, Ma et al. · Hong Kong University of Science and Technology·26 min·Jun 15, 2026
  39. 144
    When an AI Agent Just Copies Its Tool — And Bigger Models Copy More
    When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More
    Wang, Vemuri · raptorX.ai·15 min·Jun 15, 2026
  40. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
    Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack
    Hoang, Le, Xu et al. · Singapore University of Technology and Design·23 min·Jun 05, 2026
  41. 070
    When Models Know the Answer But Say the Wrong Thing Anyway
    Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer
    Yeom, Sok, Kim et al. · Graduate School of Data Science·22 min·May 22, 2026
  42. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
    Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most
    Merrill, Lee, Karger · Forecasting Research Institute / UC Berkeley·30 min·May 22, 2026
  43. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
    The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
    Liu, Holz, Ye et al. · University of Chinese Academy of Sciences·32 min·May 19, 2026
  44. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
    Training on Documents About Monitoring Leads to CoT Obfuscation
    Haskins, Chughtai, Engels · University of Canterbury·26 min·May 18, 2026
  45. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
    Exploration Hacking: Can LLMs Learn to Resist RL Training?
    Jang, Falck, Braun et al. · MATS·23 min·May 02, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.