Glossary · Term

SWE-bench

← all terms

Definition

Plain language

A benchmark that tests AI agents on real bug-fixing tasks pulled from open-source GitHub projects.

As stated in the literature

A benchmark of real-world GitHub issues from popular Python projects paired with held-out tests, used to evaluate autonomous coding agents end-to-end.

Also called: SWE-bench Verified

Why it matters: It tests coding agents on the messy, real-world software work they're being marketed for, rather than on toy programming puzzles.

For example, an agent might be given a real Django bug report and asked to produce a patch that passes the maintainers' actual test suite.

Heard on the show

“The headline number is just over 79 percent on a benchmark called SWE-bench Verified — which, Finn, we should ground for anyone who hasn't bumped into it before.”
Episode 012 — Why AI Coding Agents Keep Trying to Debug Without a Debugger

Related concepts

Related terms