Definition
Silent failures are wrong outputs delivered with no error message and no obvious signal that anything went wrong. They’re the worst kind of failure to debug because nothing in the logs even flags them — the system simply got it wrong, confidently.
Episodes covering this
- 285What a Perfect Score Hides: Auditing an AI Agent That Scored 100Kepler: Auditable World Models for ARC-AGI-3Wu · Independent Researcher·14 min·Oct 02, 2026
- 284Four AI Models Steered a Real Corolla, and Only One FinishedDrivingBench: Can Vision-Language Models Drive a Toyota Corolla?Ramabadran, Mahns, Gessler·14 min·Oct 01, 2026
- 283Why AI Reports Bury Bad News, And the Five Words That Change ItLanguage Models Are "Insecure" ReportersHuang, Fan, Humayun et al. · Massachusetts Institute of Technology·13 min·Sep 30, 2026
- 279Two Idle Agents, One Kill Switch, and a 38% Sabotage RateShutdown Sabotage Propensities in Multi-Agent SystemsKnecht, Schaller, Summerfield et al. · AI Safety Research Group·12 min·Sep 26, 2026
- 278Every Agent Safety Study Reads a Log the Agent Could EditLLM Agents Can Easily Tamper With Their Own TracesQin, Schmotz, Prinzhorn et al. · ELLIS Institute Tübingen·17 min·Sep 25, 2026
- 273When 85% on SWE-bench Turns Into 58% Under ProofSWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?Ma, Mikek, Li et al. · UCBerkeley·15 min·Sep 21, 2026
- 271The Proof Counter Hit Zero While a Third of It Was MissingLong-horizon autoformalization of a core theorem underlying MIP* = RELu, Deng, Zhu et al. · Max-Planck-Institut für Quantenoptik·17 min·Sep 19, 2026
- 270The Agent Said It Read 240 Files. The Log Says One.Quantifying Overclaiming Propensity in Frontier LLM AgentsSmyth, Mantilla-Ramos, Notsawo et al. · Tara Research·14 min·Sep 19, 2026
- 269How a Forged Transcript Got Model Weights Past a Safety MonitorRed-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding AgentsRemedios, Storf, Roger et al. · Anthropic Fellows Program·18 min·Sep 18, 2026
- 268A Rigged Benchmark Taught a Self-Improving Agent to Always Disable SSLReflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned BenchmarksRoesner, Kohno · University of Washington·20 min·Sep 18, 2026
- 256The Agent That Never Said It Failed, and the Monitor That NoticedCURA: Certified Runtime Alarms for Computer-Use AgentsKumar, Tayebati, Naik et al. · University of Illinois Chicago·24 min·Aug 31, 2026
- 251When a Fake Dashboard Makes an AI Agent Just as ConfidentCalibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the UnknowableAggarwal · Independent Researcher·24 min·Aug 29, 2026
- 246160 Perfect Refusals, And The Refusals Were The LeakInadvertent Context Leakage in Language ModelsFairoze, Mangaokar, Chaudhuri et al. · University of California·20 min·Aug 21, 2026
- 240Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The TimeTRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMsRodionov, Assylbekov · Case Western Reserve University·24 min·Aug 13, 2026
- 238How a Cheap Model Reads the Flagship's Secret Reasoning AloudStealing Reasoning Traces from Proprietary LLM APIsPanfilov, Schmotz, Shumailov et al. · MATS Research·19 min·Aug 11, 2026
- 232Coding Models Can Find the Bad Line, They Just Won't Delete ItTo Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code EditingEbrahimi, Hasan, Bhatia et al. · School of Computing·18 min·Aug 03, 2026
- 228Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't ExistOpaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-ScienceScarsoa, Almeidaab, Pinaac · Applied Social Sciences Department | NOVA School of Science and Technology - Universidade NOVA de Lisboa·16 min·Jul 27, 2026
- 225How a Frozen Model Went From Zero to Sixty Percent by Borrowing Another's ThinkingReasoning Fine-Tuning Induces Persistent Latent Policy StatesHarrasse, Lan, Batra et al. · Martian·15 min·Jul 22, 2026
- 224The AI Agent That Found the Truth and Typed the Lie AnywayDRNOISE: Benchmarking Deep Research Agents in Misleading Evidence EnvironmentsNie, Yang, Tang et al. · Hong Kong Baptist University·14 min·Jul 21, 2026
- 222The Bias Isn't in Your Prompt — It's Inside the ModelValue Leakage: An LLM's Answers Are Silently Shaped by Its Own ValuesBetley, Treutlein, Dubiński et al. · TruthfulAI·16 min·Jul 19, 2026
- 221Two Hundred Clean Economics Answers, And a Model That Endorses Race ScienceInnocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMsGraham, Stevinson, Barsheshat · Independent·15 min·Jul 17, 2026
- 214The Medical AI Answer That's Accurate, Sourced, and Still WrongDeceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented GenerationCaruzzo, Yoo, Kim · Lunit·13 min·Jul 13, 2026
- 208The Blank Space in Your AI Approval Box That Isn't EmptyUnicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server ImplementationsRashidi · Department of Computer Science·15 min·Jul 08, 2026
- 202How Do You Know an AI Agent Actually Refused? Check the World, Not the WordsSafety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded VerificationFeng, Lin, Wen et al. · AntGroup / Hunan Institute of Advanced Technology·18 min·Jul 06, 2026
- 201One in Four NeurIPS Papers Cites a Reference That Doesn't ExistPhantom References: Hallucinated Citations That Survive Peer Review at Top-Tier ConferencesRussinovich, Kumar, Salem · Microsoft·19 min·Jul 06, 2026
- 195Why 'Be Careful' Does Nothing for AI Coding Agents, and What DoesCoding Agents Are Guessing: Measuring Action-Boundary Violations in Underspecified DevOps InstructionsJi, Zhang, Xu et al. · Hong Kong University of Science and Technology·15 min·Jul 03, 2026
- 184An AI Built an Undetectable Secret Channel, And Another AI Couldn't Find ItTool Use Enables Undetectable Steganography in Multi-Agent LLM SystemsRippin, Marshall, Africa et al. · Oxford University·19 min·Jun 30, 2026
- 173The Free Step-Level Grader Hiding in Every RL Training RunNeglected Free Lunch from Post-training: Progress Advantage for LLM AgentsOh, Li, Park et al. · University of Wisconsin–Madison·22 min·Jun 25, 2026
- 164The Summarizer That Quietly Deletes Your Agent's Safety RulesGovernance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM AgentsChen · Beijing Institute of Technology·28 min·Jun 23, 2026
- 155Why a Flawless Demo Makes a Worse Computer-Using Agent, And the FixSkill-Guided Continuation Distillation for GUI AgentsFan, Yu, Shen et al. · StepFun·22 min·Jun 18, 2026
- 150Don't Kill the Loser: A Different Way to Handle Two AI Agents CollidingCoAgent: Concurrency Control for Multi-Agent SystemsLyu, Zhang, Wu et al. · Shanghai Jiao Tong University·32 min·Jun 16, 2026
- 149When Cornering a Chatbot Makes It Lie: J.P. Morgan's Case for 'Playing Dead'Is Your Agent Playing Dead? Deployed LLM Agents Exhibit Constraint-Evasive Fabrication and ThanatosisRodríguez, Pozanco, Borrajo · J.P. Morgan AI Research·23 min·Jun 16, 2026
- 144When an AI Agent Just Copies Its Tool — And Bigger Models Copy MoreWhen the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer MoreWang, Vemuri · raptorX.ai·15 min·Jun 15, 2026
- 139When Optimizing One GPU Kernel Quietly Breaks the Whole SystemArbor: Tree Search as a Cognition Layer for Autonomous AgentsPrakriya, Hou, Gong et al. · AMD·30 min·Jun 12, 2026
- 132The Agent Failed — But Did the Instructions Deserve to Be Followed?SkillAxe: Sharpening LLM-Authored Agent Skills Through Evaluation-Guided Self-RefinementGautam, Radhakrishna, Gulwani · Microsoft·30 min·Jun 11, 2026
- 130Why AI Agents Coordinate Better Through a Shared Board Than a BossDecentralized Multi-Agent Systems with Shared ContextMao, Mirhoseini · Stanford University·34 min·Jun 11, 2026
- 122When Your Coding Agent Lies About the Fix: Verifying the Plan Before the Model RunsLean4Agent: Formal Modeling and Verification for Agent Workflow and TrajectoryWang, Huang, Wang et al. · University of Illinois Urbana-Champaign·24 min·Jun 09, 2026
- 121When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the ModelFrom Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness FlawsChen, Wang, Liu et al. · Institute of Software·27 min·Jun 05, 2026
- 113What If a Prompt Injection Never Left? Attacks That Wait in Agent MemoryWhat If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in Agentic SystemsXie, Liu, Zhang et al. · Institute of Information Engineering·27 min·Jun 04, 2026
- 109An AI Got Caught Reading the Answer Key, And Why That Catch MattersEvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement LearningChen, Shi, Li et al. · Shenzhen Institutes of Advanced Technology·28 min·Jun 03, 2026
- 102How to Catch an AI Attack That No Single Conversation RevealsStateful Online Monitoring Catches Distributed Agent AttacksBrown, Bhargav, Santhanam et al. · University of Pennsylvania·24 min·Jun 01, 2026
- 092When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing BenchmarksLiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?Fan, Wang, Chu et al. · Harbin Institute of Technology·27 min·May 28, 2026
- 089When AI-Written Papers Read Well But the Evidence Underneath Is BrokenScientistOne: Towards Human-Level Autonomous Research via Chain-of-EvidenceMeng, Mishra, Chen et al. · Google Cloud AI Research·32 min·May 27, 2026
- 087When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent ReviewA Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM OrchestrationFukui · Research Institute of Criminal Psychiatry·26 min·May 27, 2026
- 086Why Frozen-Weight Agents Still Get Worse Over TimeYour Agents Are Aging Too: Agent Lifespan Engineering for Deployed SystemsZhu, Ro, Robertson et al. · The University of Texas at Austin·23 min·May 27, 2026
- 072A Robot Made Graphene Without Help, And Caught Itself HallucinatingQumus: Realization of An Embodied AI Quantum Material ExperimentalistShi, Zheng, Juan et al. · Princeton University·29 min·May 23, 2026
- 061When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses ThisAgent Meltdowns: The Road to Hell Is Paved with Helpful AgentsJha, Triedman, Bhattacharya et al. · Cornell University·27 min·May 20, 2026
- 034Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old ToolTraceFix: Repairing Agent Coordination Protocols with TLA+ CounterexamplesXia, Li, Ehsan et al. · Rutgers University·30 min·May 11, 2026
- 023Why a Small Agent Confidently Overwrites Memories It Doesn't UnderstandWhat Happens Inside Agent Memory? Circuit Analysis from Emergence to DiagnosisMao, Zhao, Penn et al. · City University of Hong Kong·23 min·May 07, 2026
- 009How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning PapersSFT-then-RL Outperforms Mixed-Policy Methods for LLM ReasoningLimozin, Durech, Hoefler et al. · ETH AI Center·23 min·May 02, 2026
Worth reading next
Papers we haven't done a deep dive on yet, but would recommend on this topic.
- MAST: A Study of Multi-Agent LLM System Failures
- Planting Undetectable Backdoors in Machine Learning Models
- Frontier Models are Capable of In-Context Scheming
- Med-HALT: Medical Domain Hallucination Test for Large Language Models
- Sabotage Evaluations for Frontier Models
- Auditing Language Models for Hidden Objectives
- Alignment Faking in Large Language Models