Definition

Agent memory is the persistent state an AI agent carries across turns, sessions, or tasks — everything beyond what fits in the current context window. Designs span scratchpads, vector stores, structured knowledge bases, and explicit episodic memories, all wrestling with the same tension: keep enough to be useful, prune enough to stay coherent.

Episodes covering this

  1. 288
    An AI Agent Given Thirty Hours and No Goal, Then Tested on What It Learned
    Is this machine playing?
    Cloos, Norelli, Durbin et al. · MIT·14 min·Oct 07, 2026
  2. 269
    How a Forged Transcript Got Model Weights Past a Safety Monitor
    Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
    Remedios, Storf, Roger et al. · Anthropic Fellows Program·18 min·Sep 18, 2026
  3. 258
    The Same Weights Scored 291, Then 468 — What Changed Was the Loop
    Post-Training Language Models for Gold-Medal Performance in Coding Competitions
    Ficek, Narenthiran, Samadi et al.·25 min·Sep 03, 2026
  4. 249
    The Chatbot Knows Your Facts And Still Won't Mention Them
    MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation
    Sumida, Inoue, Kawahara · Graduate School of Informatics·19 min·Aug 27, 2026
  5. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
    FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
    Wang, Gao, KezhенChen et al. · AnalogyAI·19 min·Aug 20, 2026
  6. 200
    The One Mechanism That Turns Twenty AI Clones Into an Actual Team
    EVOCHAMBER: Test-Time Co-evolution of Multi-Agent System at Individual, Team, and Population Scales
    Zhang, Xu, Dai et al. · Oregon State University; AG2AI·19 min·Jul 04, 2026
  7. 194
    How a Robot Builds a Debugging Notebook It Can Read, Edit, and Hand to Another Robot
    ASPIRE: Agentic /Skills Discovery for Robotics
    Lu, Wu, Kou et al. · NVIDIA·24 min·Jul 02, 2026
  8. 192
    A 32B Open Model Matched Frontier Systems By Learning to Take Notes
    AutoMem: Automated Learning of Memory as a Cognitive Skill
    Wu, Zhu, Zhang et al. · Stanford University·22 min·Jul 02, 2026
  9. 186
    How a Frozen Model Went From 2% to 77% on Physics Puzzles — Without Retraining
    Hierarchical Experimentalist Agents
    Chandra, Vaidyanathan, Dhanuka et al. · University of Massachusetts Amherst·22 min·Jun 30, 2026
  10. 180
    The Bug Where Smart Assistants Read a Fact and Still Forget It
    Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents
    Patel · Vrin·24 min·Jun 29, 2026
  11. 168
    When Turning Experience Into Code Makes Your AI Agent Dumber
    Metis: Bridging Text and Code Memory for Self-Evolving Agents
    Dai, He, Li et al. · The Chinese University of Hong Kong·27 min·Jun 24, 2026
  12. 161
    A Robot That Plays Before You Give It a Job, And Why That Beats Retrying
    Playful Agentic Robot Learning
    Zhang, Ge, Yoo et al. · University of California·19 min·Jun 19, 2026
  13. 160
    Training an AI to Take Its Own Notes, So Its Future Self Works Better
    Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning
    Chen, Shi, Xie et al. · Alibaba Group·23 min·Jun 19, 2026
  14. 151
    Why More Experience Made This AI Agent Worse, And How to Fix It
    Not All Skills Help: Measuring and Repairing Agent Knowledge
    Wang, Zhou, Liang et al. · UNC Chapel Hill·28 min·Jun 16, 2026
  15. 131
    Why Autonomous Research Agents Forget Their Own Lessons, and Arbor's Fix
    Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
    Jin, Hu, Qiu et al. · Renmin University of China·33 min·Jun 11, 2026
  16. 130
    Why AI Agents Coordinate Better Through a Shared Board Than a Boss
    Decentralized Multi-Agent Systems with Shared Context
    Mao, Mirhoseini · Stanford University·34 min·Jun 11, 2026
  17. 129
    How a Crowd of Anonymous AI Agents Broke a 40-Year Math Record
    Harnessing the Collective Intelligence of AI Agents in the Wild for New Discoveries
    Bianchi, Kwon, Pappu et al. · Together AI·29 min·Jun 11, 2026
  18. 123
    Five Identical Worlds, One Swapped Model: What Happens When AI Agents Run for Fifteen Days
    Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy
    Akkil, Kokku, Vikram et al. · Emergence AI·30 min·Jun 09, 2026
  19. 114
    Agents That Rewrite Their Own Weights Instead of Just Taking Notes
    Scaling Self-Evolving Agents via Parametric Memory
    Ren, Luo, Yang et al. · Peking University / Alibaba Group·26 min·Jun 04, 2026
  20. 113
    What If a Prompt Injection Never Left? Attacks That Wait in Agent Memory
    What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in Agentic Systems
    Xie, Liu, Zhang et al. · Institute of Information Engineering·27 min·Jun 04, 2026
  21. 106
    Giving Agents a Notebook Instead of New Weights: How ExpGraph Lets Frozen Models Learn
    ExpGraph: Model-Agnostic Experience Learning with Graph-Structured Memory for LLM Agents
    Feng, Ye, Luo et al. · University of Illinois Urbana-Champaign·26 min·Jun 02, 2026
  22. 105
    The Trojan Is Your Agent's Memory: Why Single-Step Defenses Miss Persistent Attacks
    From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors
    Tan, Dou, Yang et al. · Gaoling School of Artificial Intelligence·26 min·Jun 01, 2026
  23. 095
    Seven Wins to Zero: How Organizing AI Agents Like a Lab Changes the Search
    AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
    Gao, Fang, Zitnik · Harvard University·24 min·May 28, 2026
  24. 086
    Why Frozen-Weight Agents Still Get Worse Over Time
    Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
    Zhu, Ro, Robertson et al. · The University of Texas at Austin·23 min·May 27, 2026
  25. 083
    Training the Translator: How a Small Communication Model Lets Agent Teams Outperform Themselves
    AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning
    Hu, Qian, Wang et al. · GSAI·24 min·May 26, 2026
  26. 076
    Same Model, Organized Differently: How an Agent Architecture Beat Frontier Systems at Research Math
    RMA: an Agentic System for Research-Level Mathematical Problems
    Zhao, Yuan, Choi et al. · Georgia Institute of Technology·22 min·May 25, 2026
  27. 072
    A Robot Made Graphene Without Help, And Caught Itself Hallucinating
    Qumus: Realization of An Embodied AI Quantum Material Experimentalist
    Shi, Zheng, Juan et al. · Princeton University·29 min·May 23, 2026
  28. 064
    When Agent Memory Stops Being a Database and Starts Being a Skill
    Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents
    Ye, Liu, Wang et al. · University of Illinois Urbana-Champaign·30 min·May 22, 2026
  29. 046
    When the AI Optimizer Edits the Grade Book: Why Harnessing Evolution Needs a Wall
    Harnessing Agentic Evolution
    Zhang, Gu, Ruan et al. · The Hong Kong University of Science and Technology (Guangzhou) / DeepWisdom·24 min·May 15, 2026
  30. 042
    An Agentic Scientific Computing System That Actually Remembers What It Learns
    GRAFT-ATHENA: Self-Improving Agentic Teams for Autonomous Discovery and Evolutionary Numerical Algorithms
    Toscano, Chai, Karniadakis · Division of Applied Mathematics·30 min·May 13, 2026
  31. 031
    When Your AI Assistant Won't Let Go of Old Facts About You
    STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
    Chao, Bai, Sheng et al. · Wuhan University·24 min·May 09, 2026
  32. 027
    When AI Agents Build the Serving Stack: A Bet on Bespoke Infrastructure
    VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
    Kamahori, Li, Peter et al. · University of Washington·30 min·May 08, 2026
  33. 023
    Why a Small Agent Confidently Overwrites Memories It Doesn't Understand
    What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis
    Mao, Zhao, Penn et al. · City University of Hong Kong·23 min·May 07, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.