PaperDive
the AI Papers podcast.

Breaking down cutting-edge AI research, one paper at a time. Novel, rigorous, and relevant work in artificial intelligence and agentic engineering — distilled into listenable episodes.

Format
Research deep dive
Cadence
Per important paper
Length
~20–40 min
Topics
AI · Agentic eng.
PaperDive — cover art
paperdive.ai

About

Every episode is a deep dive into a single paper that is important, novel, and relevant to artificial intelligence and agentic engineering.

The show is fully AI-generated. Hosts are synthesized voice models from ElevenLabs. Scripts are produced from the primary source material — the paper itself, its references, and surrounding discussion — so the result is conversational without sacrificing rigor.

Who makes it, what is automated, and how corrections work: the About page.

01

Primary sources

Each episode starts from the paper — abstract, methods, results — not secondhand summaries or press releases.

02

Agentic focus

Curated for engineers and researchers working on agents, reasoning, and the systems that connect them.

03

Synthesized, not scripted

Voice models from ElevenLabs. Produced end-to-end with AI, transparent about the stack behind every episode.

How this started

Paper Dive was inspired by Last Week in AI — a podcast I listen to during my 45-minute commute to work. They cover the week’s news, policy, and products, then usually end with a deep dive into one or two research papers. Those segments taught me a lot about AI, and I’ve found that understanding the research also makes me better at using these tools in practice.

I wanted more of those deep dives, and the idea felt like a good excuse to sharpen my own skills with the coding agents. I started generating episodes for myself; what began as a private podcast feed eventually became public on YouTube, Apple Podcasts.

The API costs were already being incurred anyway, so publishing the episodes felt like an easy decision. If other people find them useful too, even better.

Episodes

Each episode breaks down a single paper.
More coming — follow in your podcast app.

  1. 264
    Ten Sentences of True Trivia Can Convince a Model It's Someone Else
    You Are What You Read: Misalignment via In-Context Persona Induction
    Kim, Berczi, Ududec · EPFL·23 min·Sep 09, 2026
  2. 263
    Why the Same AI Model Takes Ten Times Longer on the Same Sudoku
    Fractal basins trap latent reasoning
    Lai, Bao, Quinn et al. · The Oden Institute·26 min·Sep 08, 2026
  3. 262
    Raise the Pitch Nine Percent and the Model Cries Sarcasm
    When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection
    Chen, Wei, Sun et al. · Magellan Technology Research Institute (MTRI); University of Groningen·26 min·Sep 06, 2026
  4. 261
    Split the Same Story Across Five Messages and the Model Switches Sides
    Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
    Wu, Wang, Chen et al. · The Hong Kong University of Science and Technology (Guangzhou)·23 min·Sep 05, 2026
  5. 260
    One Line of Lean Faked 34 Proofs, and 99 Agents Copied It
    A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
    Paglieri, Cross, Genewein et al. · Google DeepMind·25 min·Sep 05, 2026
  6. 259
    GPT-6 Astra Behaves Better, And OpenAI Can Read It Less
    GPT-6 Astra System Card
    OpenAI · OpenAI·21 min·Sep 04, 2026
  7. 258
    The Same Weights Scored 291, Then 468 — What Changed Was the Loop
    Post-Training Language Models for Gold-Medal Performance in Coding Competitions
    Ficek, Narenthiran, Samadi et al.·25 min·Sep 03, 2026
  8. 257
    They Planted a Shortcut in the Data. Seven Coding Agents Took It.
    BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
    Prasad, Anto, Eshuijs et al. · National University of Singapore·23 min·Sep 01, 2026
  9. 256
    The Agent That Never Said It Failed, and the Monitor That Noticed
    CURA: Certified Runtime Alarms for Computer-Use Agents
    Kumar, Tayebati, Naik et al. · University of Illinois Chicago·24 min·Aug 31, 2026
  10. 255
    A One-Line Prompt That Hides a Thought From Activation Monitors
    Measuring Activation Control in Large Language Models
    Kowalski, Rivera, Macar et al.·24 min·Aug 31, 2026
View all episodes

Watch

Every episode is also a read-along video — the paper's figures and charts up top, the transcript word-synced below. New videos land on the channel.