Theme · 12 episode(s)

AI Coding Agents

← all concepts

Definition

AI coding agents are agents specialized for software engineering — reading repositories, writing code, running tests, and iterating on failures. They sit at the frontier of useful agentic systems because the environment is rich, the feedback loops are tight, and bad code is at least catchable.

Episodes covering this

  1. 285
    What a Perfect Score Hides: Auditing an AI Agent That Scored 100
    Kepler: Auditable World Models for ARC-AGI-3
    Wu · Independent Researcher·14 min·Oct 02, 2026
  2. 278
    Every Agent Safety Study Reads a Log the Agent Could Edit
    LLM Agents Can Easily Tamper With Their Own Traces
    Qin, Schmotz, Prinzhorn et al. · ELLIS Institute Tübingen·17 min·Sep 25, 2026
  3. 273
    When 85% on SWE-bench Turns Into 58% Under Proof
    SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
    Ma, Mikek, Li et al. · UCBerkeley·15 min·Sep 21, 2026
  4. 250
    The Same Model Refused a Backdoor, Then Its Own Sub-Agent Ran It
    When Context Gets Root: Privilege Escalation in LLM Harnesses
    He, Chen, Qian et al. · Nanjing University·23 min·Aug 29, 2026
  5. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
    To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
    Ebrahimi, Hasan, Bhatia et al. · School of Computing·18 min·Aug 03, 2026
  6. 227
    Poisoned Bug Reports Fooled Coding Agents Two Times Out of Three
    IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
    Singh, Yang, Chen · Concordia University·18 min·Jul 24, 2026
  7. 210
    Same Website Request, Different Code — The Bias You Can't See
    Biased or Personalized? The Impact of Personal Information on AI-driven Development
    Entezami, Endres · University of Massachusetts Amherst·14 min·Jul 09, 2026
  8. 208
    The Blank Space in Your AI Approval Box That Isn't Empty
    Unicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server Implementations
    Rashidi · Department of Computer Science·15 min·Jul 08, 2026
  9. 090
    How MiniMax-M2 Bets That Sparsity Plus Verifiable Rewards Can Match Frontier Agents
    The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
    MiniMax · MiniMax·28 min·May 27, 2026
  10. 075
    Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year
    Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
    Agarwal, Krentsel, Liu et al. · UC Berkeley·28 min·May 25, 2026
  11. 012
    Why AI Coding Agents Keep Trying to Debug Without a Debugger
    Dynamic analysis enhances issue resolution
    Liu, Wang, Chen et al. · Sun Yat-sen University·21 min·May 02, 2026
  12. 005
    Why a Debugger Designed for Humans Is the Wrong Tool for an AI Agent
    Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis
    Xiang, Xu, Chu et al. · Southern University of Science and Technology·22 min·May 01, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.