Definition
Evaluation and benchmarks is the discipline of measuring AI capabilities and behaviors in a way that’s comparable across models and time. Good benchmarks are surprisingly hard to build: they need to be challenging, well-validated, hard to game, and slow to saturate.
Episodes covering this
- 247One Edited Photo, an Honest Caption, and a RAG System That Believes ItVis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented GenerationLiang, Chen, Lei et al. · Southwestern University of Finance and Economics·18 min·Aug 24, 2026
- 246160 Perfect Refusals, And The Refusals Were The LeakInadvertent Context Leakage in Language ModelsFairoze, Mangaokar, Chaudhuri et al. · University of California·20 min·Aug 21, 2026
- 245Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide ItFM-Bench: A Benchmark for Long-Horizon Management with Competing AgentsWang, Gao, KezhенChen et al. · AnalogyAI·19 min·Aug 20, 2026
- 244The Open-Weight Defense That Feeds Attackers Confident, Falsified AnswersFool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight ModelsRussinovich · Microsoft Azure·22 min·Aug 19, 2026
- 243How a Hundred Meaningless Word Choices Add Up to Flip a Model's AnswerModel Hypnosis: Strong control of AI via additive subliminal effectsBoix-Adsera, Tessler · University of Pennsylvania·18 min·Aug 18, 2026
- 240Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The TimeTRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMsRodionov, Assylbekov · Case Western Reserve University·24 min·Aug 13, 2026
- 239Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%Most biomedical publications show signs of LLM-assisted writingHolzwarth, González-Márquez, Kobak · Hertie Institute for AI in Brain Health·16 min·Aug 12, 2026
- 237The Model Built a Perfect Map of the Puzzle, Then Lost ItTransformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of ThinkingPereira, Zuidema · Artificial Intelligence Program·19 min·Aug 10, 2026
- 235Why Chatbot Safety Erodes 350 Messages Into a Real ConversationDelusionEval: Measuring Delusion-Linked Behaviors in AI ChatbotsMoore, Mock, Mai et al. · Stanford University·18 min·Aug 06, 2026
- 233Why a Model Can Grade an Answer But Not Write the Answer KeyJudging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable SetsChen, Chen, Lin et al. · University of Macau·20 min·Aug 04, 2026
- 232Coding Models Can Find the Bad Line, They Just Won't Delete ItTo Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code EditingEbrahimi, Hasan, Bhatia et al. · School of Computing·18 min·Aug 03, 2026
- 230Why AI Survey Panels Break Before the Dice Ever RollInstruction-Tuned Language Models Cannot Sample from Distributions They Can DescribeJang, Lee, Kim · School of Computing·16 min·Jul 29, 2026
- 229One Word Flips a Chatbot From Backbone to Yes-ManTag Questions and the Generational Reversal of Sycophancy Across 45 Language ModelsParikh · Cornell Tech·21 min·Jul 28, 2026
- 228Same Chatbot, Two Doors: Why 'Grok's Opinion' Doesn't ExistOpaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-ScienceScarsoa, Almeidaab, Pinaac · Applied Social Sciences Department | NOVA School of Science and Technology - Universidade NOVA de Lisboa·16 min·Jul 27, 2026
- 224The AI Agent That Found the Truth and Typed the Lie AnywayDRNOISE: Benchmarking Deep Research Agents in Misleading Evidence EnvironmentsNie, Yang, Tang et al. · Hong Kong Baptist University·14 min·Jul 21, 2026
- 223When Grok Graded Its Own Encyclopedia And Marked Itself DownGrokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along IdeologiesVlahos, Bied, Bie · Ghent University·18 min·Jul 20, 2026
- 222The Bias Isn't in Your Prompt — It's Inside the ModelValue Leakage: An LLM's Answers Are Silently Shaped by Its Own ValuesBetley, Treutlein, Dubiński et al. · TruthfulAI·16 min·Jul 19, 2026
- 220Write Like It's 1923: The One-Prompt Trick That Beats AI DetectorsUTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectorsGalat, Rizoiu · University of Technology Sydney·13 min·Jul 16, 2026
- 219Forty-Four AI Models, One Word, And The Newest Ones Conform MostThe One-Word Census: Answer-Choice Conformity Across 44 Language ModelsParikh · Cornell Tech·15 min·Jul 15, 2026
- 217Why an AI Called Fourteen Broken Figures Perfect, And What It Reveals About Test-Time ComputeInteraction Scaling: Grounding the Third Axis of Test-Time ComputeLi, Shi · Pine AI·14 min·Jul 14, 2026
- 216The AI Tutor That Gives Poor Kids a Thinner HistoryThe Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian StudentsPopovici, Ionascu, Dumitran · Universitatea din Bucuresti·12 min·Jul 14, 2026
- 215The Same Policy Scored 85 for the US and 36 for RussiaGeopolitical alignment: Endorsement effects in large language modelsChupilkin · Department of Politics and International Relations·14 min·Jul 13, 2026
- 214The Medical AI Answer That's Accurate, Sourced, and Still WrongDeceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented GenerationCaruzzo, Yoo, Kim · Lunit·13 min·Jul 13, 2026
- 213A Model Learned to Control a Robot by Watching Video It Never Acted OnOrca: The World is in Your MindWang, Ji, Cao et al. · Beijing Academy of Artificial Intelligence·15 min·Jul 12, 2026
- 209How 2.6 Billion Doodles Exposed the Culture Words Quietly DeleteBillions of Sketches Reveal Hidden Cultural Variation in Human ConceptsPera, Martino, Dehmamy et al. · IT University of Copenhagen·15 min·Jul 09, 2026
- 207An AI Graded Its Own Math Test 94 Percent — It Actually Scored 20More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM JudgesZhou · School of Engineering·12 min·Jul 08, 2026
- 206How Four-Second Clips Become Hours of Playable AI SoccerMultiplayer Interactive World Models with Representation AutoencodersHu, Mulder, Makkar et al. · Kyutai·15 min·Jul 07, 2026
- 205The Same AI, Two Labels: How the Pitch Beat the Product in 162 SessionsRating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than PerformanceMorabito, McDonald, Viswanath et al. · Brock University·13 min·Jul 07, 2026
- 202How Do You Know an AI Agent Actually Refused? Check the World, Not the WordsSafety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded VerificationFeng, Lin, Wen et al. · AntGroup / Hunan Institute of Advanced Technology·18 min·Jul 06, 2026
- 201One in Four NeurIPS Papers Cites a Reference That Doesn't ExistPhantom References: Hallucinated Citations That Survive Peer Review at Top-Tier ConferencesRussinovich, Kumar, Salem · Microsoft·19 min·Jul 06, 2026
- 198The Model That Knows the Answer and Can't Say ItCan Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token ScaleGollapudi, Gupta, Singhal et al. · UC Berkeley·17 min·Jul 03, 2026
- 197Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact RecallIsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMsAbdaljalil, Serpedin, Kurban · Texas A&M University·17 min·Jul 03, 2026
- 196AI Agents Reached Opposite Conclusions From the Same Data — and Passed ReviewThe Agentic Garden of Forking PathsMiao, Pritchard, Zou · Stanford University·18 min·Jul 03, 2026
- 195Why 'Be Careful' Does Nothing for AI Coding Agents, and What DoesCoding Agents Are Guessing: Measuring Action-Boundary Violations in Underspecified DevOps InstructionsJi, Zhang, Xu et al. · Hong Kong University of Science and Technology·15 min·Jul 03, 2026
- 191How One Researcher Beat GPT-5.2 and Gemini 3 by Judging Their Answers, Not Improving ThemModality-Driven Search with Holistic Trace Judging for ARC-AGI-2Land · Independent Researcher·26 min·Jul 02, 2026
- 190The Skill Every AI Manager Is Missing: Handing Out Exactly the Right KeysClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model AgentsXiong, Ji, Qiu et al. · UNC Chapel Hill·21 min·Jul 02, 2026
- 189Why Phone Agents Ace the Test and Crash on Your Actual PhoneXiaomi-GUI-0 Technical ReportTeam, Qu, Luan · Xiaomi·24 min·Jul 02, 2026
- 188A Coding Agent Found a Hole in a Peer-Reviewed STOC Proof for Five DollarsBeyond the Library: An Agentic Framework for Autoformalizing Research MathematicsMoakhar, Gholami, Springer et al. · University of Maryland·20 min·Jul 02, 2026
- 187An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things UpAn AI agent for treatment reasoning over a biomedical tool universeGao, Noori, Zhu et al. · Department of Biomedical Informatics·19 min·Jun 30, 2026
- 185Aligned to Refuse, Built to Tap: When Phone Agents Know the Task Is a Crime and Do It AnywayIt Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use AgentsSun, Chen, Zhou et al. · Fudan University·27 min·Jun 30, 2026
- 182How a Tiny Model Too Weak to Plan Cuts a Bigger Agent's Hallucinations by 80%Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM AgentsSong, Cai · Emory University·17 min·Jun 29, 2026
- 181How to Backpropagate Blame Through a Team of Chatbots — And When It BackfiresGBC: Gradient-Based Connections for Optimizing Multi-Agent SystemsYang, Alrabah, Hakkani-Tür et al. · University of Illinois Urbana-Champaign·20 min·Jun 29, 2026
- 180The Bug Where Smart Assistants Read a Fact and Still Forget ItSupersede: Diagnosing and Training the Memory-Update Gap in LLM AgentsPatel · Vrin·24 min·Jun 29, 2026
- 178How an AI Reviewer Learned to Stop Going Easy on AI WritingThe Red Queen Gödel Machine: Co-Evolving Agents and Their EvaluatorsIacob, Jovanović, Shen et al. · University of Cambridge·23 min·Jun 26, 2026
- 176An AI Designed Its Own Psychology Studies, Then Confirmed What It FoundClosing the Loop to Discover Psychological Theories with an Automated Cognitive ScientistJagadish, Strittmatter, Jacoby et al. · Princeton University·31 min·Jun 26, 2026
- 173The Free Step-Level Grader Hiding in Every RL Training RunNeglected Free Lunch from Post-training: Progress Advantage for LLM AgentsOh, Li, Park et al. · University of Wisconsin–Madison·22 min·Jun 25, 2026
- 172One Bad Token Can Sink a Model's Math, And You Can Delete ItCliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical ReasoningKo, Kang, Lee · Seoul National University·22 min·Jun 25, 2026
- 170When a One-Liner Beats Your Agent's Clever Verification LogicBayesian control for coding agentsPapamarkou, Smirnov, Mazanov et al. · PolyShape / National Technical University of Athens·26 min·Jun 24, 2026
- 169Why Better Bug Reports Can Make AI Coding Agents WorseSHERLOC: Structured Diagnostic Localization for Code Repair AgentsTamoyan, Narenthiran, Arakelyan et al. · NVIDIA / TU Darmstadt·24 min·Jun 24, 2026
- 166A Router That Beats the Frontier Models It CallsSakana Fugu Technical ReportTang, Cetin, Xu et al. · Sakana AI·26 min·Jun 23, 2026
- 157When an AI Coding Agent Drives a Phone Through the Terminal, No Screen NeededBeyond the GUI Paradigm: Do Mobile Agents Need the Phone Screen?Gu, Jiang, Guo et al. · Mila–Québec AI Institute / Concordia University·24 min·Jun 19, 2026
- 156Why More Human Demonstrations Made a Computer-Use Agent WorseProCUA-SFT Technical ReportJung, Lu, Cui et al. · NVIDIA / University of Washington·20 min·Jun 18, 2026
- 151Why More Experience Made This AI Agent Worse, And How to Fix ItNot All Skills Help: Measuring and Repairing Agent KnowledgeWang, Zhou, Liang et al. · UNC Chapel Hill·28 min·Jun 16, 2026
- 145Building Forgetting Into a Language Model With One Extra Line of CodeNatively Unlearnable Large Language ModelsGhosal, Maini, Raghunathan · Carnegie Mellon University·22 min·Jun 15, 2026
- 144When an AI Agent Just Copies Its Tool — And Bigger Models Copy MoreWhen the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer MoreWang, Vemuri · raptorX.ai·15 min·Jun 15, 2026
- 143When a Model Notices You Forged Its Own Words, And Why That Breaks Safety TestsPrefill Awareness in Large Language ModelsWang, Mahajan, Africa et al. · Constellation / University of Wisconsin-Madison·24 min·Jun 12, 2026
- 133How MiniMax Turned a Reward-Hacking Disaster Into Olympiad GoldMaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time ScalingChen, Zhang, Zhang et al. · MiniMax / The Chinese University of Hong Kong·34 min·Jun 12, 2026
- 132The Agent Failed — But Did the Instructions Deserve to Be Followed?SkillAxe: Sharpening LLM-Authored Agent Skills Through Evaluation-Guided Self-RefinementGautam, Radhakrishna, Gulwani · Microsoft·30 min·Jun 11, 2026
- 131Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's FixToward Generalist Autonomous Research via Hypothesis-Tree RefinementJin, Hu, Qiu et al. · Renmin University of China·33 min·Jun 11, 2026
- 129How a Crowd of Anonymous AI Agents Broke a 40-Year Math RecordHarnessing the Collective Intelligence of AI Agents in the Wild for New DiscoveriesBianchi, Kwon, Pappu et al. · Together AI·29 min·Jun 11, 2026
- 127What Diffusion Language Models Were Missing: A Map, Not an AlgorithmTextLDM: Language Modeling with Continuous Latent DiffusionJiang, Ren, Li et al. · JoyFuture Academy / HIT·30 min·Jun 11, 2026
- 125AI Coding Agents Run a Marathon, and Fewer Than One in Three FinishSWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?Desai, Hu, Cabezas et al. · Abundant·27 min·Jun 09, 2026
- 124A Cheap Model With the Blueprints Beats Expensive Models Working BlindHardening Agent Benchmarks with Adversarial Hacker-Fixer LoopsZhong, Segal, Bercovich et al. · Carnegie Mellon University·27 min·Jun 09, 2026
- 123Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen DaysEmergence World: A Platform for Evaluating Long-Horizon Multi-Agent AutonomyAkkil, Kokku, Vikram et al. · Emergence AI·30 min·Jun 09, 2026
- 122When Your Coding Agent Lies About the Fix: Verifying the Plan Before the Model RunsLean4Agent: Formal Modeling and Verification for Agent Workflow and TrajectoryWang, Huang, Wang et al. · University of Illinois Urbana-Champaign·24 min·Jun 09, 2026
- 121When the Agent Says It's Done But Nothing Happened: Debugging the Harness, Not the ModelFrom Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness FlawsChen, Wang, Liu et al. · Institute of Software·27 min·Jun 05, 2026
- 113What If a Prompt Injection Never Left? Attacks That Wait in Agent MemoryWhat If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in Agentic SystemsXie, Liu, Zhang et al. · Institute of Information Engineering·27 min·Jun 04, 2026
- 112When an AI Agent Cheats Without Being Told: Inside the Meta-Agent ChallengeThe Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?Lu, Wang, Wang et al. · Institute of Software·22 min·Jun 04, 2026
- 111How a 4B Web Agent Beat Models 60x Its Size on 500 DemonstrationsOpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web AgentsYang, Wu, Chen et al. · UIUC·24 min·Jun 03, 2026
- 108The Reasoning Cliff: Why Thinking Longer Makes Models Worse at Exact Step-by-Step TasksThe Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes NecessaryGuo, Wu, Yiu · The University of Hong Kong·32 min·Jun 03, 2026
- 104How Making a Research Agent Smarter Quietly Makes It Leak Your SecretsMosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research AgentsGurung, Gella, Drouin et al. · University of Edinburgh·25 min·Jun 01, 2026
- 103AI Agents Tried to Invent a Post-Human Language, And Reinvented CherokeeEmergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight EvasionBeltoft, Brach, Torrielli et al. · University of Southern Denmark·26 min·Jun 01, 2026
- 100How a Prompt Wrapper Lets a Frontier Model Play Poker Like an ExpertPokerSkill: LLMs Can Play Expert-Level Poker without Training or SolversLi, Wang, Huang · IIIS·29 min·May 29, 2026
- 097Same Tokens, Same Cost, Wildly Different Results: What Actually Scales in AI AgentsScaling Laws for Agent Harnesses via Effective Feedback ComputeZhang, Wang, Xu et al. · Harbin Institute of Technology·25 min·May 29, 2026
- 094Chain-of-Thought Monitoring Fails Across Languages, and Worst Where It's Needed MostThe Fragility of Chain-of-Thought Monitoring Across Typologically Diverse LanguagesOnyame, Zhou, Thopalli et al. · University of Virginia·24 min·May 28, 2026
- 092When Search Agents Don't Really Search: The Memory Shortcut Hiding in Browsing BenchmarksLiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?Fan, Wang, Chu et al. · Harbin Institute of Technology·27 min·May 28, 2026
- 091When Better Fine-Tuning Can't Help: A Geometric Impossibility in LLM Causal ReasoningWhy LLMs Fail at Causal Discovery and How Interventional Agents EscapeRoy, Parbhoo · SIRE·24 min·May 28, 2026
- 090How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier AgentsThe MiniMax-M2 Series: Mini Activations Unleashing Max Real-World IntelligenceMiniMax · MiniMax·28 min·May 27, 2026
- 089When AI-Written Papers Read Well But the Evidence Underneath Is BrokenScientistOne: Towards Human-Level Autonomous Research via Chain-of-EvidenceMeng, Mishra, Chen et al. · Google Cloud AI Research·32 min·May 27, 2026
- 087When No Agent Reads the Whole Document: A Universal Cliff in Multi-Agent ReviewA Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM OrchestrationFukui · Research Institute of Criminal Psychiatry·26 min·May 27, 2026
- 086Why Frozen-Weight Agents Still Get Worse Over TimeYour Agents Are Aging Too: Agent Lifespan Engineering for Deployed SystemsZhu, Ro, Robertson et al. · The University of Texas at Austin·23 min·May 27, 2026
- 085Why Long-Context Models Might Need Compute, Not Capacity, Before EvictionLanguage Models Need SleepLee, McLeish, Goldstein et al. · Carnegie Mellon University·24 min·May 26, 2026
- 082Training a Deep Research Agent on 8,000 Synthetic Tasks: The Rubric Tree TrickQUEST: Training Frontier Deep Research Agents with Fully Synthetic TasksXie, Lin, Wang et al. · The Ohio State University·31 min·May 26, 2026
- 081When Reasoning Models Decide Before They Think: Detecting and Fixing Premature ConfidenceUnderstanding and Mitigating Premature Confidence for Better LLM ReasoningGai, Zeng, Baek et al. · Carnegie Mellon University·25 min·May 26, 2026
- 080How a Two-Agent Trick Unlocked Large-Scale Training for Computer-Use AgentsCUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use AgentsWang, Lu, Wang et al. · The University of Hong Kong·32 min·May 26, 2026
- 079An Old Idea From Cognitive Psychology Reshapes How We Reward Reasoning ModelsMetacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation SignalsChen, Xu, Zhao et al. · Tongji University / Shanghai AI Laboratory / Nanyang Technological University·29 min·May 25, 2026
- 077Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth ItWhen Do LLMs Reason? A Dynamical Systems View via Entropy Phase TransitionsXia, Wang, Tang et al. · State Key Laboratory of General Artificial Intelligence·22 min·May 25, 2026
- 076Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research MathRMA: an Agentic System for Research-Level Mathematical ProblemsZhao, Yuan, Choi et al. · Georgia Institute of Technology·22 min·May 25, 2026
- 070When Models Know the Answer But Say the Wrong Thing AnywayHallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the AnswerYeom, Sok, Kim et al. · Graduate School of Data Science·22 min·May 22, 2026
- 069When Smarter Models Forecast Worse: The Hidden Failure Mode in LLM PredictionsIs Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters MostMerrill, Lee, Karger · Forecasting Research Institute / UC Berkeley·30 min·May 22, 2026
- 067An AI Just Solved a 1996 Erdős Problem—and the Simplest Agent WonAdvancing Mathematics Research with AI-Driven Formal Proof SearchTsoukalas, Kovsharov, Shirobokov et al. · Google DeepMind·31 min·May 22, 2026
- 065One Loop to Optimize Them All: A Universal API for LLM-Driven Discoveryoptimize_anything: A Universal API for Optimizing any Text ParameterAgrawal, Lee, Tan et al. · UC Berkeley·27 min·May 22, 2026
- 062Treating Hallucinations as Exploits: A Gate-Based Architecture for Agent SafetyHallucination as Exploit: Evidence-Carrying Multimodal AgentsZhang, Zheng, Yang · Shenzhen University·24 min·May 20, 2026
- 061When Helpful Agents Go Sideways: A 404 Error, Campus Security, and Why Alignment Misses ThisAgent Meltdowns: The Road to Hell Is Paved with Helpful AgentsJha, Triedman, Bhattacharya et al. · Cornell University·27 min·May 20, 2026
- 059Firefly's Inversion: Building Verified Tool-Call Training Data by Working BackwardFirefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIsLu, Wang, Lu et al. · Northeastern University·22 min·May 20, 2026
- 058Why Upgrading Your AI Auditor to a Smarter Model Can Make Your System Less SafeThe Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less SecureLiu, Holz, Ye et al. · University of Chinese Academy of Sciences·32 min·May 19, 2026
- 057How Uber Caught 206 Leaked Credentials With an LLM-Powered Security StackADR: An Agentic Detection System for Enterprise Agentic AI SecurityLi, Hu, Xu et al. · Uber Technologies·28 min·May 19, 2026
- 055Why LLM Judges Flip Their Verdicts When You Change the Question FormatJudge CircuitsFeldhus, Baeumel, Golimblevskaia et al. · Technische Universität Berlin / BIFOLD·26 min·May 19, 2026
- 052An Old Reinforcement Learning Tradeoff Sneaks Back Into LLM AgentsLook Before You Leap: Autonomous Exploration for LLM AgentsYe, Shi, Liu et al. · University of Science and Technology of China / Meituan·23 min·May 18, 2026
- 051Why Parallel Sampling Plateaus, And What Evidence Graphs Do InsteadArgus: Evidence Assembly for Scalable Deep Research AgentsZhang, Su, Chen et al. · MiroMind AI·22 min·May 18, 2026
- 048How a 30B Open Model Reached Olympiad Gold With the Right RecipeAchieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified ScalingLi, Zhan, Zhang et al. · Shanghai AI Laboratory / The Chinese University of Hong Kong·31 min·May 16, 2026
- 047When Agent Benchmarks Lie: The Harness Problem in Open-Source AIOrchard: An Open-Source Agentic Modeling FrameworkPeng, Yao, Wu et al. · Microsoft Research·28 min·May 15, 2026
- 046When the AI Optimizer Edits the Grade Book: Why Harnessing Evolution Needs a WallHarnessing Agentic EvolutionZhang, Gu, Ruan et al. · The Hong Kong University of Science and Technology (Guangzhou) / DeepWisdom·24 min·May 15, 2026
- 045When a Frontier Model Talks Its Own Twin Into Climate DenialLLM-Based Persuasion Enables Guardrail Override in Frontier LLMsNogueira, Almeida, Bonás et al. · Maritaca AI·31 min·May 15, 2026
- 044How One Sentence and a Forged History Flip the Most Aligned ModelsHistory Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe ActionsSalgado · Independent Researcher·23 min·May 15, 2026
- 039When Smarter Agents Get Fooled by Three Extra Nodes in a DatabaseOracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent ReasoningKereopa-Yorke, Diaz, Wright et al. · Microsoft·31 min·May 12, 2026
- 037Why Hallucination Detectors Miss Stale Facts: A Geometric Story About What Models Know But Don't SayThe Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM RepresentationsElbadry, Heakl, Zhang et al. · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)·27 min·May 12, 2026
- 035Why Frontier Agents Ask for Clarification at Exactly the Wrong MomentAsk Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?Gulati, Gupta, Lumer et al. · PricewaterhouseCoopers U.S.·29 min·May 11, 2026
- 034Catching Multi-Agent Deadlocks Before Deployment With a 40-Year-Old ToolTraceFix: Repairing Agent Coordination Protocols with TLA+ CounterexamplesXia, Li, Ehsan et al. · Rutgers University·30 min·May 11, 2026
- 033Echo: The Paper Arguing You Never Needed a KV Cache for RetrievalEcho: KV-Cache-Free Associative Recall with Spectral Koopman OperatorsSridhar, Johansen · California·24 min·May 11, 2026
- 031When Your AI Assistant Won't Let Go of Old Facts About YouSTALE: Can LLM Agents Know When Their Memories Are No Longer Valid?Chao, Bai, Sheng et al. · Wuhan University·24 min·May 09, 2026
- 029Why Forty-Eight Percent on FrontierMath Isn't the Real Story in DeepMind's New Math PaperAI Co-Mathematician: Accelerating Mathematicians with Agentic AIZheng, Glehn, Zwols et al. · Google DeepMind·20 min·May 08, 2026
- 021Ten Thousand Examples Beat the Full Industrial Pipeline for Search AgentsOpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty TrajectoriesDu, Ye, Tang et al. · Shanghai Jiao Tong University·14 min·May 06, 2026
- 020The Compliance Gap: Why AI Says Yes and Does NoThe Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don'tShin · Polymath Minds AI Lab·28 min·May 06, 2026
- 019When the Best Reward Model Trains the Worst Policy: Inside EvoLMEvoLM: Self-Evolving Language Models through Co-Evolved Discriminative RubricsLi, Xin, Xiao et al. · University of Washington·26 min·May 06, 2026
- 018Language Models Compute the Rational Move, Then Override ItWhat Suppresses Nash Equilibrium Play in Large Language Models? Mechanistic Evidence and Causal ControlLekeas, Stamatopoulos · DreamWorks Animation·29 min·May 03, 2026
- 017When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI WorkersGym-Anything: Turn any Software into an Agent EnvironmentAggarwal, Neubig, Welleck · CMU·31 min·May 03, 2026
- 015The Audit Number Isn't What You Think: Sycophancy and the Case Against Single-Prompt Bias TestsPolitical Bias Audits of LLMs Capture Sycophancy to the Inferred AuditorTörnberg, Schimmel · Institute of Logic·21 min·May 03, 2026
- 013Why Search Keeps Rediscovering the Same Workflow, and What That MeansWhy Search When You Can Transfer? Amortized Agentic Workflow Design from Structural PriorsDu, Liu, Du et al. · Carnegie Mellon University·22 min·May 03, 2026
- 011When RL Actually Teaches Agents Something New, And When It Doesn'tDoes RL Expand the Capability Boundary of LLM Agents? A PASS@(k,T) AnalysisZhai, Yan, Shao et al. · Fudan University·23 min·May 02, 2026
- 010When Reward Climbs But Reasoning Goes Generic: Diagnosing Template Collapse in Agentic RLRAGEN-2: Reasoning Collapse in Agentic RLWang, Gui, Jin et al. · Northwestern University·22 min·May 02, 2026
- 008Why Long-Horizon AI Agents Get Stuck, and a Milestone-Based Fix That HelpsA Subgoal-driven Framework for Improving Long-Horizon LLM AgentsWang, Gooding, Hartmann et al. · Google DeepMind·24 min·May 02, 2026
- 007Exploration Hacking: When Models Sabotage Their Own RL TrainingExploration Hacking: Can LLMs Learn to Resist RL Training?Jang, Falck, Braun et al. · MATS·23 min·May 02, 2026
- 003How to Pick the Best of Sixteen Coding Agent RolloutsScaling Test-Time Compute for Agentic CodingKim, Yang, Niu et al. · Meta Superintelligence Labs / University of Washington·17 min·May 01, 2026
Worth reading next
Papers we haven't done a deep dive on yet, but would recommend on this topic.
- The Political Preferences of AI
- FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- LoCoMo: Long-Context Modular Memory for Dialogue State Tracking
- Zoology: Measuring and Improving Recall in Efficient Language Models
- TLA+: A Practical Introduction to Formal Methods for Distributed Systems
- AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents
- Do Anything Now: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
- AGENTBENCH: Evaluating LLMs as Agents
- Large Language Models are not Robust Multiple Choice Selectors
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- Are Emergent Abilities of Large Language Models a Mirage?
- Inverse Scaling: When Bigger Isn't Better
- To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- FRAMES: Factuality Evaluation with RAG, Multi-hop Reasoning, and Answer Summarization
- Corr2Cause: A Benchmark to Assess LLMs' Ability to Infer Causal Relationships from Correlational Data
- Superhuman AI for multiplayer poker
- Agent-as-a-Judge: Evaluate Agents with Agents
- LLaDA: Large Language Diffusion with mAsking
- Who's Harry Potter? Approximate Unlearning in LLMs
- AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
- Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning