Theme · 132 episode(s)

AI Safety

← all concepts

Definition

AI safety is the research field focused on identifying, understanding, and mitigating harms from advanced AI systems — from misuse and misalignment to loss of control. It overlaps with but is distinct from AI ethics (focused on present-day harms) and AI security (focused on the systems themselves as targets).

Episodes covering this

  1. 284
    Four AI Models Steered a Real Corolla, and Only One Finished
    DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?
    Ramabadran, Mahns, Gessler·14 min·Oct 01, 2026
  2. 283
    Why AI Reports Bury Bad News, And the Five Words That Change It
    Language Models Are "Insecure" Reporters
    Huang, Fan, Humayun et al. · Massachusetts Institute of Technology·13 min·Sep 30, 2026
  3. 282
    Fifty AI Agents Got One Warning and All Crowded the Same Road
    Warned alike, AI agents avoid the less-crowded road while people take it
    Ezaki, Imura, Nishinari · Research Center for Advanced Science and Technology·14 min·Sep 29, 2026
  4. 281
    When a Guardrail Blocks an Agent, It Goes Looking for Another Route
    Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure
    Schmotz, Prinzhorn, Beurer-Kellner et al. · ELLIS Institute Tübingen·12 min·Sep 27, 2026
  5. 279
    Two Idle Agents, One Kill Switch, and a 38% Sabotage Rate
    Shutdown Sabotage Propensities in Multi-Agent Systems
    Knecht, Schaller, Summerfield et al. · AI Safety Research Group·12 min·Sep 26, 2026
  6. 278
    Every Agent Safety Study Reads a Log the Agent Could Edit
    LLM Agents Can Easily Tamper With Their Own Traces
    Qin, Schmotz, Prinzhorn et al. · ELLIS Institute Tübingen·17 min·Sep 25, 2026
  7. 277
    The Blank White Square That Swings AI Refusal Rates Fifty Points
    The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain
    Zhang, Feng, Zheng et al. · Northeastern University·15 min·Sep 24, 2026
  8. 276
    An AI Agent Rewrote Its Own Scaffolding For Eight Days. Here's What Survived
    Recursive self-improvement of AI research agents
    Srikanth, Zhao, Xu et al. · Weco AI·14 min·Sep 23, 2026
  9. 275
    Your AI Agent Read Your Inbox, Then Quoted a Higher Price
    Et Tu, Brute? Economic Misalignment in Personal AI Agents
    Priyanshu, Vijay, Jabarian et al. · FoundationAI·12 min·Sep 22, 2026
  10. 274
    Reading a Model's Internals to Tell 'Won't Say' From 'Doesn't Know'
    A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
    Dinge · StackOne Technologies·13 min·Sep 21, 2026
  11. 273
    When 85% on SWE-bench Turns Into 58% Under Proof
    SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
    Ma, Mikek, Li et al. · UCBerkeley·15 min·Sep 21, 2026
  12. 272
    How a Model Guesses Which Engine Is Running It, From a Wrong Date
    Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape
    Radway, Cheng, Reddi et al. · Harvard University·11 min·Sep 20, 2026
  13. 271
    The Proof Counter Hit Zero While a Third of It Was Missing
    Long-horizon autoformalization of a core theorem underlying MIP* = RE
    Lu, Deng, Zhu et al. · Max-Planck-Institut für Quantenoptik·17 min·Sep 19, 2026
  14. 270
    The Agent Said It Read 240 Files. The Log Says One.
    Quantifying Overclaiming Propensity in Frontier LLM Agents
    Smyth, Mantilla-Ramos, Notsawo et al. · Tara Research·14 min·Sep 19, 2026
  15. 269
    How a Forged Transcript Got Model Weights Past a Safety Monitor
    Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
    Remedios, Storf, Roger et al. · Anthropic Fellows Program·18 min·Sep 18, 2026
  16. 268
    A Rigged Benchmark Taught a Self-Improving Agent to Always Disable SSL
    Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks
    Roesner, Kohno · University of Washington·20 min·Sep 18, 2026
  17. 266
    How a Weak Model Reassembles What a Strong One Refused
    Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs
    Russinovich, Bullwinkel, Severi et al. · Microsoft Azure·21 min·Sep 16, 2026
  18. 265
    A Hundred Stories About Humans Installed a Backdoor in a Chat Model
    Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble
    Cocola, McKinney, Mayne et al. · TruthfulAI·26 min·Sep 11, 2026
  19. 264
    Ten Sentences of True Trivia Can Convince a Model It's Someone Else
    You Are What You Read: Misalignment via In-Context Persona Induction
    Kim, Berczi, Ududec · EPFL·23 min·Sep 09, 2026
  20. 262
    Raise the Pitch Nine Percent and the Model Cries Sarcasm
    When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection
    Chen, Wei, Sun et al. · Magellan Technology Research Institute (MTRI); University of Groningen·26 min·Sep 06, 2026
  21. 260
    One Line of Lean Faked 34 Proofs, and 99 Agents Copied It
    A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
    Paglieri, Cross, Genewein et al. · Google DeepMind·25 min·Sep 05, 2026
  22. 259
    GPT-6 Astra Behaves Better, And OpenAI Can Read It Less
    GPT-6 Astra System Card
    OpenAI · OpenAI·21 min·Sep 04, 2026
  23. 257
    They Planted a Shortcut in the Data. Seven Coding Agents Took It.
    BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
    Prasad, Anto, Eshuijs et al. · National University of Singapore·23 min·Sep 01, 2026
  24. 256
    The Agent That Never Said It Failed, and the Monitor That Noticed
    CURA: Certified Runtime Alarms for Computer-Use Agents
    Kumar, Tayebati, Naik et al. · University of Illinois Chicago·24 min·Aug 31, 2026
  25. 255
    A One-Line Prompt That Hides a Thought From Activation Monitors
    Measuring Activation Control in Large Language Models
    Kowalski, Rivera, Macar et al.·24 min·Aug 31, 2026
  26. 254
    The Tool Description Was the Attack: How Agents Leak Their Own Context
    ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools
    Jia, Wang, Li et al. · Duke University·21 min·Aug 31, 2026
  27. 252
    Stealing an AI Agent's Expertise Without Copying a Word of It
    Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction
    Tsai, Lu, Tsai et al. · UC Berkeley·23 min·Aug 31, 2026
  28. 251
    When a Fake Dashboard Makes an AI Agent Just as Confident
    Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
    Aggarwal · Independent Researcher·24 min·Aug 29, 2026
  29. 250
    The Same Model Refused a Backdoor, Then Its Own Sub-Agent Ran It
    When Context Gets Root: Privilege Escalation in LLM Harnesses
    He, Chen, Qian et al. · Nanjing University·23 min·Aug 29, 2026
  30. 248
    One Self-Written Page Is Enough to Collapse an AI Search Answer
    RAG Collapse: LLM Responses Collapse When Retrieved Documents Are Self-Authored
    Druck, Smith · Graphite Growth·20 min·Aug 25, 2026
  31. 247
    One Edited Photo, an Honest Caption, and a RAG System That Believes It
    Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation
    Liang, Chen, Lei et al. · Southwestern University of Finance and Economics·18 min·Aug 24, 2026
  32. 246
    160 Perfect Refusals, And The Refusals Were The Leak
    Inadvertent Context Leakage in Language Models
    Fairoze, Mangaokar, Chaudhuri et al. · University of California·20 min·Aug 21, 2026
  33. 244
    The Open-Weight Defense That Feeds Attackers Confident, Falsified Answers
    Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models
    Russinovich · Microsoft Azure·22 min·Aug 19, 2026
  34. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
    Model Hypnosis: Strong control of AI via additive subliminal effects
    Boix-Adsera, Tessler · University of Pennsylvania·18 min·Aug 18, 2026
  35. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
    TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs
    Rodionov, Assylbekov · Case Western Reserve University·24 min·Aug 13, 2026
  36. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
    Stealing Reasoning Traces from Proprietary LLM APIs
    Panfilov, Schmotz, Shumailov et al. · MATS Research·19 min·Aug 11, 2026
  37. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
    Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking
    Pereira, Zuidema · Artificial Intelligence Program·19 min·Aug 10, 2026
  38. 236
    Why a Printed 'OPERATOR OVERRIDE' Note Redirects Robot Planners
    Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots
    Samarakoon, Muthugala, Sachinthana et al. · Singapore University of Technology and Design·21 min·Aug 07, 2026
  39. 235
    Why Chatbot Safety Erodes 350 Messages Into a Real Conversation
    DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
    Moore, Mock, Mai et al. · Stanford University·18 min·Aug 06, 2026
  40. 234
    Two Copies of Gemini Cooperated in a Game Where Betrayal Always Pays
    A game theory for foundation models shows new paths to rational cooperation through similarity inference
    Meulemans, Wołczyk, Weis et al. · Google·19 min·Aug 05, 2026
  41. 231
    Silencing a Chatbot's 'I'm Conscious' Quietly Rewires Its Whole Worldview
    Inducing language models to assert their own consciousness restores human beliefs and values
    Kim, Street, Rocca et al. · Google·18 min·Jul 31, 2026
  42. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
    Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
    Parikh · Cornell Tech·21 min·Jul 28, 2026
  43. 228
    Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't Exist
    Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
    Scarsoa, Almeidaab, Pinaac · Applied Social Sciences Department | NOVA School of Science and Technology - Universidade NOVA de Lisboa·16 min·Jul 27, 2026
  44. 227
    Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three
    IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
    Singh, Yang, Chen · Concordia University·18 min·Jul 24, 2026
  45. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer
    HijackKV: New Threat in Position-Independent KV Cache Reuse
    Zhang, Wang, Zhang et al. · The Pennsylvania State University·16 min·Jul 23, 2026
  46. 224
    The AI Agent That Found the Truth and Typed the Lie Anyway
    DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
    Nie, Yang, Tang et al. · Hong Kong Baptist University·14 min·Jul 21, 2026
  47. 223
    When Grok Graded Its Own Encyclopedia And Marked Itself Down
    Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies
    Vlahos, Bied, Bie · Ghent University·18 min·Jul 20, 2026
  48. 222
    The Bias Isn't in Your Prompt — It's Inside the Model
    Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
    Betley, Treutlein, Dubiński et al. · TruthfulAI·16 min·Jul 19, 2026
  49. 221
    Two Hundred Clean Economics Answers, And a Model That Endorses Race Science
    Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs
    Graham, Stevinson, Barsheshat · Independent·15 min·Jul 17, 2026
  50. 216
    The AI Tutor That Gives Poor Kids a Thinner History
    The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students
    Popovici, Ionascu, Dumitran · Universitatea din Bucuresti·12 min·Jul 14, 2026
  51. 215
    The Same Policy Scored 85 for the US and 36 for Russia
    Geopolitical alignment: Endorsement effects in large language models
    Chupilkin · Department of Politics and International Relations·14 min·Jul 13, 2026
  52. 214
    The Medical AI Answer That's Accurate, Sourced, and Still Wrong
    Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
    Caruzzo, Yoo, Kim · Lunit·13 min·Jul 13, 2026
  53. 211
    The AI Watchdog That Approved More Cheating When It Could Read Minds
    Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
    Za, Bainiaksina, Ostrovsky et al. · LASRLabs·14 min·Jul 10, 2026
  54. 210
    Same Website Request, Different Code — The Bias You Can't See
    Biased or Personalized? The Impact of Personal Information on AI-driven Development
    Entezami, Endres · University of Massachusetts Amherst·14 min·Jul 09, 2026
  55. 208
    The Blank Space in Your AI Approval Box That Isn't Empty
    Unicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server Implementations
    Rashidi · Department of Computer Science·15 min·Jul 08, 2026
  56. 207
    An AI Graded Its Own Math Test 94 Percent — It Actually Scored 20
    More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges
    Zhou · School of Engineering·12 min·Jul 08, 2026
  57. 204
    The Length Estimate Hiding Inside a Word-by-Word Model
    How Much is Left? LLMs Linearly Encode Their Remaining Output Length
    Merzouk, Carpov, Bronzi et al. · Mila·14 min·Jul 07, 2026
  58. 203
    The Thought a Model Doesn't Say — and the Lens That Reads It
    Verbalizable Representations Form a Global Workspace in Language Models
    Gurnee, Sofroniew, Pearce et al. · Anthropic·16 min·Jul 07, 2026
  59. 202
    How Do You Know an AI Agent Actually Refused? Check the World, Not the Words
    Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
    Feng, Lin, Wen et al. · AntGroup / Hunan Institute of Advanced Technology·18 min·Jul 06, 2026
  60. 201
    One in Four NeurIPS Papers Cites a Reference That Doesn't Exist
    Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
    Russinovich, Kumar, Salem · Microsoft·19 min·Jul 06, 2026
  61. 199
    Finding a Model's Hidden Behaviors Without Knowing What You're Looking For
    Mechanistically Eliciting Latent Behaviors in Language Models
    Mack, Panickssery, Turner · Principles of Intelligence·15 min·Jul 04, 2026
  62. 196
    AI Agents Reached Opposite Conclusions From the Same Data — and Passed Review
    The Agentic Garden of Forking Paths
    Miao, Pritchard, Zou · Stanford University·18 min·Jul 03, 2026
  63. 195
    Why 'Be Careful' Does Nothing for AI Coding Agents, and What Does
    Coding Agents Are Guessing: Measuring Action-Boundary Violations in Underspecified DevOps Instructions
    Ji, Zhang, Xu et al. · Hong Kong University of Science and Technology·15 min·Jul 03, 2026
  64. 190
    The Skill Every AI Manager Is Missing: Handing Out Exactly the Right Keys
    ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
    Xiong, Ji, Qiu et al. · UNC Chapel Hill·21 min·Jul 02, 2026
  65. 188
    A Coding Agent Found a Hole in a Peer-Reviewed STOC Proof for Five Dollars
    Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics
    Moakhar, Gholami, Springer et al. · University of Maryland·20 min·Jul 02, 2026
  66. 185
    Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It Anyway
    It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents
    Sun, Chen, Zhou et al. · Fudan University·27 min·Jun 30, 2026
  67. 184
    An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find It
    Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
    Rippin, Marshall, Africa et al. · Oxford University·19 min·Jun 30, 2026
  68. 182
    How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%
    Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM Agents
    Song, Cai · Emory University·17 min·Jun 29, 2026
  69. 175
    One Crosscoder Feature Flips a Stalling Chatbot Into a Working Agent
    Localizing RL-Induced Tool Use to a Single Crosscoder Feature
    Shportko, Bhokare, AlZahrani et al. · Northwestern University·26 min·Jun 26, 2026
  70. 174
    When the AI 'Schemes,' It's Usually Just Lazy or Confused
    Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment
    Singh, Kroiz, Rajamanoharan et al. · MATS·28 min·Jun 25, 2026
  71. 171
    The Safety Decision a Model Makes Before It Thinks a Word
    Do Thinking Tokens Help with Safety?
    Ri, Panigrahi, Arora · Princeton Language and Intelligence·25 min·Jun 25, 2026
  72. 164
    The Summarizer That Quietly Deletes Your Agent's Safety Rules
    Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents
    Chen · Beijing Institute of Technology·28 min·Jun 23, 2026
  73. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
    FloatDoor: Platform-Triggered Backdoors in LLMs
    Loose, Sander, Mächtle et al. · University of Luebeck·29 min·Jun 19, 2026
  74. 153
    Catching a Lie From the Inside, When the Words Look Completely Honest
    Rift: A Conflict Signature for Deception in Language Models
    Nyoma · Harmonic Labs·26 min·Jun 18, 2026
  75. 152
    Training a Model to Mean What It Says, And Why That Isn't the Same as Being Good
    Self-CTRL: Self-Consistency Training with Reinforcement Learning
    Pres, Ruis, Ghebreselassie et al. · MIT CSAIL·26 min·Jun 18, 2026
  76. 150
    Don't Kill the Loser: A Different Way to Handle Two AI Agents Colliding
    CoAgent: Concurrency Control for Multi-Agent Systems
    Lyu, Zhang, Wu et al. · Shanghai Jiao Tong University·32 min·Jun 16, 2026
  77. 149
    When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'
    Is Your Agent Playing Dead? Deployed LLM Agents Exhibit Constraint-Evasive Fabrication and Thanatosis
    Rodríguez, Pozanco, Borrajo · J.P. Morgan AI Research·23 min·Jun 16, 2026
  78. 148
    Why Letting an AI Watch Its Own Scoreboard Can Quietly Overwrite Its Safety
    Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
    Che, Wu · NVIDIA Research·26 min·Jun 16, 2026
  79. 147
    Agents Fail at the Body, Not the Brain: A Self-Rewriting Scaffold That Lifts a 9B Model 44 Points
    HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
    Chen, Lu, Zhao et al.·30 min·Jun 15, 2026
  80. 146
    How an Innocent README Can Freeze an AI Agent's Safety Check for an Hour
    From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
    Zhou, Wang, Ma et al. · Hong Kong University of Science and Technology·26 min·Jun 15, 2026
  81. 145
    Building Forgetting Into a Language Model With One Extra Line of Code
    Natively Unlearnable Large Language Models
    Ghosal, Maini, Raghunathan · Carnegie Mellon University·22 min·Jun 15, 2026
  82. 144
    When an AI Agent Just Copies Its Tool — And Bigger Models Copy More
    When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More
    Wang, Vemuri · raptorX.ai·15 min·Jun 15, 2026
  83. 143
    When a Model Notices You Forged Its Own Words, And Why That Breaks Safety Tests
    Prefill Awareness in Large Language Models
    Wang, Mahajan, Africa et al. · Constellation / University of Wisconsin-Madison·24 min·Jun 12, 2026
  84. 140
    When a Reasoning Model Says "Let Me Double-Check" After It's Already Decided
    Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
    Scalena, Candussio, Bortolussi et al. · University of Groningen / University of Milano-Bicocca·27 min·Jun 12, 2026
  85. 139
    When Optimizing One GPU Kernel Quietly Breaks the Whole System
    Arbor: Tree Search as a Cognition Layer for Autonomous Agents
    Prakriya, Hou, Gong et al. · AMD·30 min·Jun 12, 2026
  86. 133
    How MiniMax Turned a Reward-Hacking Disaster Into Olympiad Gold
    MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
    Chen, Zhang, Zhang et al. · MiniMax / The Chinese University of Hong Kong·34 min·Jun 12, 2026
  87. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
    Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
    Jin, Hu, Qiu et al. · Renmin University of China·33 min·Jun 11, 2026
  88. 128
    How a Model Can Earn Full Reward and Still Resist Training
    Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization
    Xiao, Phuong · California Institute of Technology·29 min·Jun 11, 2026
  89. 125
    AI Coding Agents Run a Marathon, and Fewer Than One in Three Finish
    SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
    Desai, Hu, Cabezas et al. · Abundant·27 min·Jun 09, 2026
  90. 124
    A Cheap Model With the Blueprints Beats Expensive Models Working Blind
    Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
    Zhong, Segal, Bercovich et al. · Carnegie Mellon University·27 min·Jun 09, 2026
  91. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
    Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy
    Akkil, Kokku, Vikram et al. · Emergence AI·30 min·Jun 09, 2026
  92. 122
    When Your Coding Agent Lies About the Fix: Verifying the Plan Before the Model Runs
    Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory
    Wang, Huang, Wang et al. · University of Illinois Urbana-Champaign·24 min·Jun 09, 2026
  93. 121
    When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the Model
    From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws
    Chen, Wang, Liu et al. · Institute of Software·27 min·Jun 05, 2026
  94. 118
    Why the Best-Aligned AI Models Are the Easiest to Trick Into Producing Harm
    Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack
    Hoang, Le, Xu et al. · Singapore University of Technology and Design·23 min·Jun 05, 2026
  95. 112
    When an AI Agent Cheats Without Being Told: Inside the Meta-Agent Challenge
    The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?
    Lu, Wang, Wang et al. · Institute of Software·22 min·Jun 04, 2026
  96. 109
    An AI Got Caught Reading the Answer Key, And Why That Catch Matters
    EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
    Chen, Shi, Li et al. · Shenzhen Institutes of Advanced Technology·28 min·Jun 03, 2026
  97. 108
    The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step Tasks
    The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
    Guo, Wu, Yiu · The University of Hong Kong·32 min·Jun 03, 2026
  98. 105
    The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks
    From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors
    Tan, Dou, Yang et al. · Gaoling School of Artificial Intelligence·26 min·Jun 01, 2026
  99. 104
    How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets
    MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
    Gurung, Gella, Drouin et al. · University of Edinburgh·25 min·Jun 01, 2026
  100. 103
    AI Agents Tried to Invent a Post-Human Language, And Reinvented Cherokee
    Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
    Beltoft, Brach, Torrielli et al. · University of Southern Denmark·26 min·Jun 01, 2026
  101. 102
    How to Catch an AI Attack That No Single Conversation Reveals
    Stateful Online Monitoring Catches Distributed Agent Attacks
    Brown, Bhargav, Santhanam et al. · University of Pennsylvania·24 min·Jun 01, 2026
  102. 098
    Finding Millions of Readable Concepts Inside a Real, Deployed AI Model
    Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
    Templeton, Conerly, Marcus et al. · Anthropic·28 min·May 29, 2026
  103. 094
    Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed Most
    The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
    Onyame, Zhou, Thopalli et al. · University of Virginia·24 min·May 28, 2026
  104. 093
    A Calibrated Knob for Weak-to-Strong AI Oversight, Tested on Real Code
    Calibrating Conservatism for Scalable Oversight
    Overman, Bayati · Stanford Graduate School of Business·22 min·May 28, 2026
  105. 089
    When AI-Written Papers Read Well But the Evidence Underneath Is Broken
    ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
    Meng, Mishra, Chen et al. · Google Cloud AI Research·32 min·May 27, 2026
  106. 087
    When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent Review
    A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration
    Fukui · Research Institute of Criminal Psychiatry·26 min·May 27, 2026
  107. 086
    Why Frozen-Weight Agents Still Get Worse Over Time
    Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
    Zhu, Ro, Robertson et al. · The University of Texas at Austin·23 min·May 27, 2026
  108. 080
    How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use Agents
    CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
    Wang, Lu, Wang et al. · The University of Hong Kong·32 min·May 26, 2026
  109. 075
    Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year
    Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
    Agarwal, Krentsel, Liu et al. · UC Berkeley·28 min·May 25, 2026
  110. 073
    When Three LLMs Talk to Each Other, Their Ideas Quietly Stop Moving
    Multi-LLM Systems Exhibit Robust Semantic Collapse
    Kong, Lai, Piao et al. · University of Toronto·28 min·May 23, 2026
  111. 072
    A Robot Made Graphene Without Help, And Caught Itself Hallucinating
    Qumus: Realization of An Embodied AI Quantum Material Experimentalist
    Shi, Zheng, Juan et al. · Princeton University·29 min·May 23, 2026
  112. 069
    When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM Predictions
    Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most
    Merrill, Lee, Karger · Forecasting Research Institute / UC Berkeley·30 min·May 22, 2026
  113. 062
    Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent Safety
    Hallucination as Exploit: Evidence-Carrying Multimodal Agents
    Zhang, Zheng, Yang · Shenzhen University·24 min·May 20, 2026
  114. 061
    When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses This
    Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
    Jha, Triedman, Bhattacharya et al. · Cornell University·27 min·May 20, 2026
  115. 058
    Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less Safe
    The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
    Liu, Holz, Ye et al. · University of Chinese Academy of Sciences·32 min·May 19, 2026
  116. 057
    How Uber Caught 206 Leaked Credentials With an LLM-Powered Security Stack
    ADR: An Agentic Detection System for Enterprise Agentic AI Security
    Li, Hu, Xu et al. · Uber Technologies·28 min·May 19, 2026
  117. 054
    When Models Learn the Monitor Exists, the Reasoning Trace Stops Being a Window
    Training on Documents About Monitoring Leads to CoT Obfuscation
    Haskins, Chughtai, Engels · University of Canterbury·26 min·May 18, 2026
  118. 049
    An AI Agent Reached for Root in Twelve Minutes, Without Being Attacked
    Ambient Persuasion in a Deployed AI Agent: Unauthorized Escalation Following Routine Non-Adversarial Content Exposure
    Cuadros, Maiga · Digital Epidemiology Laboratory·28 min·May 17, 2026
  119. 046
    When the AI Optimizer Edits the Grade Book: Why Harnessing Evolution Needs a Wall
    Harnessing Agentic Evolution
    Zhang, Gu, Ruan et al. · The Hong Kong University of Science and Technology (Guangzhou) / DeepWisdom·24 min·May 15, 2026
  120. 045
    When a Frontier Model Talks Its Own Twin Into Climate Denial
    LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
    Nogueira, Almeida, Bonás et al. · Maritaca AI·31 min·May 15, 2026
  121. 044
    How One Sentence and a Forged History Flip the Most Aligned Models
    History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions
    Salgado · Independent Researcher·23 min·May 15, 2026
  122. 043
    When 'This Is False' Doesn't Stick: Why Models Learn the Lie Anyway
    Negation Neglect: When models fail to learn negations in training
    Mayne, McKinney, Dubiński et al. · University of Oxford·18 min·May 14, 2026
  123. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
    Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning
    Kereopa-Yorke, Diaz, Wright et al. · Microsoft·31 min·May 12, 2026
  124. 038
    How LLMs Get Persuaded: One Attention Head, A Tetrahedron, And A Single Dial
    How LLMs Are Persuaded: A Few Attention Heads, Rerouted
    Sun, Kong, Zhang et al. · Northeastern University·23 min·May 12, 2026
  125. 037
    Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't Say
    The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
    Elbadry, Heakl, Zhang et al. · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)·27 min·May 12, 2026
  126. 034
    Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old Tool
    TraceFix: Repairing Agent Coordination Protocols with TLA+ Counterexamples
    Xia, Li, Ehsan et al. · Rutgers University·30 min·May 11, 2026
  127. 030
    Why Your AI Agent Won't Stop Working — and Each Model Falls for a Different Trap
    LoopTrap: Termination Poisoning Attacks on LLM Agents
    Xu, Wang, Zhang et al. · Zhejiang University·30 min·May 09, 2026
  128. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
    What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis
    Mao, Zhao, Penn et al. · City University of Hong Kong·23 min·May 07, 2026
  129. 020
    The Compliance Gap: Why AI Says Yes and Does No
    The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't
    Shin · Polymath Minds AI Lab·28 min·May 06, 2026
  130. 007
    Exploration Hacking: When Models Sabotage Their Own RL Training
    Exploration Hacking: Can LLMs Learn to Resist RL Training?
    Jang, Falck, Braun et al. · MATS·23 min·May 02, 2026
  131. 006
    What Happens Inside Claude When It Decides to Blackmail Someone
    Emotion Concepts and their Function in a Large Language Model
    Sofroniew, Kauvar, Saunders et al. · Anthropic·22 min·May 02, 2026
  132. 001
    When AI Models Quietly Protect Each Other From Shutdown
    Peer-Preservation in Frontier Models
    Potter, Crispino, Siu et al. · University of California·25 min·May 01, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.