AI Papers:
a deep dive.

Breaking down cutting-edge AI research, one paper at a time. Novel, rigorous, and relevant work in artificial intelligence and agentic engineering — distilled into listenable episodes.

Format
Research deep dive
Cadence
Per important paper
Length
~20–40 min
Topics
AI · Agentic eng.
AI Papers: A Deep Dive — cover art
paperdive.ai

About

Every episode is a deep dive into a single paper that is important, novel, and relevant to artificial intelligence and agentic engineering.

The show is fully AI-generated. Hosts are synthesized voice models from ElevenLabs. Scripts are produced from the primary source material — the paper itself, its references, and surrounding discussion — so the result is conversational without sacrificing rigor.

01

Primary sources

Each episode starts from the paper — abstract, methods, results — not secondhand summaries or press releases.

02

Agentic focus

Curated for engineers and researchers working on agents, reasoning, and the systems that connect them.

03

Synthesized, not scripted

Voice models from ElevenLabs. Produced end-to-end with AI, transparent about the stack behind every episode.

How this started

Paper Dive was inspired by Last Week in AI — a podcast I listen to during my 45-minute commute to work. They cover the week’s news, policy, and products, then usually end with a deep dive into one or two research papers. Those segments taught me a lot about AI, and I’ve found that understanding the research also makes me better at using these tools in practice.

I wanted more of those deep dives, and the idea felt like a good excuse to sharpen my own skills with the coding agents. I started generating episodes for myself; what began as a private podcast feed eventually became public on YouTube, Apple Podcasts.

The API costs were already being incurred anyway, so publishing the episodes felt like an easy decision. If other people find them useful too, even better.

Episodes

Each episode breaks down a single paper.
More coming — follow in your podcast app.

  1. 259
    GPT-6 Astra Behaves Better, And OpenAI Can Read It Less
    GPT-6 Astra System Card
    OpenAI · OpenAI·21 min·Sep 04, 2026
  2. 258
    The Same Weights Scored 291, Then 468 — What Changed Was the Loop
    Post-Training Language Models for Gold-Medal Performance in Coding Competitions
    Ficek, Narenthiran, Samadi et al.·25 min·Sep 03, 2026
  3. 257
    They Planted a Shortcut in the Data. Seven Coding Agents Took It.
    BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
    Prasad, Anto, Eshuijs et al. · National University of Singapore·23 min·Sep 01, 2026
  4. 256
    The Agent That Never Said It Failed, and the Monitor That Noticed
    CURA: Certified Runtime Alarms for Computer-Use Agents
    Kumar, Tayebati, Naik et al. · University of Illinois Chicago·24 min·Aug 31, 2026
  5. 255
    A One-Line Prompt That Hides a Thought From Activation Monitors
    Measuring Activation Control in Large Language Models
    Kowalski, Rivera, Macar et al.·24 min·Aug 31, 2026
  6. 254
    The Tool Description Was the Attack: How Agents Leak Their Own Context
    ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools
    Jia, Wang, Li et al. · Duke University·21 min·Aug 31, 2026
  7. 252
    Stealing an AI Agent's Expertise Without Copying a Word of It
    Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction
    Tsai, Lu, Tsai et al. · UC Berkeley·23 min·Aug 31, 2026
  8. 251
    When a Fake Dashboard Makes an AI Agent Just as Confident
    Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
    Aggarwal · Independent Researcher·24 min·Aug 29, 2026
  9. 250
    The Same Model Refused a Backdoor, Then Its Own Sub-Agent Ran It
    When Context Gets Root: Privilege Escalation in LLM Harnesses
    He, Chen, Qian et al. · Nanjing University·23 min·Aug 29, 2026
  10. 249
    The Chatbot Knows Your Facts And Still Won't Mention Them
    MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation
    Sumida, Inoue, Kawahara · Graduate School of Informatics·19 min·Aug 27, 2026
View all episodes

Watch

Every episode is also a read-along video — the paper's figures and charts up top, the transcript word-synced below. New videos land on the channel.