Theme · 7 episode(s)

LLM Agents

← all concepts

Definition

LLM agents are agents whose decision-making core is a large language model: it reads the situation, picks the next tool call, writes the next message, decides when it’s done. Almost every interesting AI agent in 2026 is an LLM agent in this sense.

Episodes covering this

  1. 238
    How a Cheap Model Reads the Flagship's Secret Reasoning Aloud
    Stealing Reasoning Traces from Proprietary LLM APIs
    Panfilov, Schmotz, Shumailov et al. · MATS Research·19 min·Aug 11, 2026
  2. 237
    The Model Built a Perfect Map of the Puzzle, Then Lost It
    Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking
    Pereira, Zuidema · Artificial Intelligence Program·19 min·Aug 10, 2026
  3. 208
    The Blank Space in Your AI Approval Box That Isn't Empty
    Unicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server Implementations
    Rashidi · Department of Computer Science·15 min·Jul 08, 2026
  4. 187
    An 8-Billion Agent That Beats Models 80 Times Its Size By Looking Things Up
    An AI agent for treatment reasoning over a biomedical tool universe
    Gao, Noori, Zhu et al. · Department of Biomedical Informatics·19 min·Jun 30, 2026
  5. 104
    How Making a Research Agent Smarter Quietly Makes It Leak Your Secrets
    MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
    Gurung, Gella, Drouin et al. · University of Edinburgh·25 min·Jun 01, 2026
  6. 028
    Teaching a Model to Hire Copies of Itself: Recursive Agent Optimization
    Recursive Agent Optimization
    Gandhi, Chakraborty, Wang et al. · Carnegie Mellon University·23 min·May 08, 2026
  7. 021
    Ten Thousand Examples Beat the Full Industrial Pipeline for Search Agents
    OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories
    Du, Ye, Tang et al. · Shanghai Jiao Tong University·14 min·May 06, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.