Theme · 7 episode(s)

Software Engineering Automation

← all concepts

Definition

Software engineering automation covers the tools and agents that take over parts of the SE workflow — writing code, reviewing diffs, running tests, triaging bugs, drafting PRs. The frontier question is which parts cleanly automate and which stubbornly require human judgment.

Episodes covering this

  1. 232
    Coding Models Can Find the Bad Line, They Just Won't Delete It
    To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
    Ebrahimi, Hasan, Bhatia et al. · School of Computing·18 min·Aug 03, 2026
  2. 169
    Why Better Bug Reports Can Make AI Coding Agents Worse
    SHERLOC: Structured Diagnostic Localization for Code Repair Agents
    Tamoyan, Narenthiran, Arakelyan et al. · NVIDIA / TU Darmstadt·24 min·Jun 24, 2026
  3. 126
    How Coding Agents Can Mine Their Own Failures Into a Self-Targeting Curriculum
    Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
    Xiao, Jiao, Wang et al. · Shanghai Jiao Tong University·21 min·Jun 09, 2026
  4. 075
    Growing Code and Proof Together: Verified Systems in Ten Hours Instead of a Year
    Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
    Agarwal, Krentsel, Liu et al. · UC Berkeley·28 min·May 25, 2026
  5. 039
    When Smarter Agents Get Fooled by Three Extra Nodes in a Database
    Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning
    Kereopa-Yorke, Diaz, Wright et al. · Microsoft·31 min·May 12, 2026
  6. 012
    Why AI Coding Agents Keep Trying to Debug Without a Debugger
    Dynamic analysis enhances issue resolution
    Liu, Wang, Chen et al. · Sun Yat-sen University·21 min·May 02, 2026
  7. 005
    Why a Debugger Designed for Humans Is the Wrong Tool for an AI Agent
    Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis
    Xiang, Xu, Chu et al. · Southern University of Science and Technology·22 min·May 01, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.