Definition
Plain language
A benchmark of hard command-line tasks for agentic systems.
As stated in the literature
A command-line task benchmark for AI agents covering operations like file recovery, system administration, and shell-driven problem solving.
Also called: TerminalBench 2.0, Terminal-Bench
Why it matters: It tests whether agents can really operate a computer the way a sysadmin does, not just write code in a sandbox.
For example, an agent might be dropped into a broken Linux system and asked to recover deleted files using only shell commands.
Heard on the show
“… On the benchmark they care about — TerminalBench 2.0, which is just a public suite of real terminal tasks that real agents struggle with — Qwen3-8B …”Episode 084 — Terminal Agents Get Free Supervision From The Tokens We've Been Throwing Away