Definition
Plain language
A benchmark of hard command-line tasks for agentic systems.
As stated in the literature
A command-line task benchmark for AI agents covering operations like file recovery, system administration, and shell-driven problem solving.
Also called: Terminal-Bench v-two
Why it matters: It tests whether agents can really operate a computer the way a sysadmin does, not just write code in a sandbox.
For example, an agent might be dropped into a broken Linux system and asked to recover deleted files using only shell commands.
Heard on the show
“On Terminal-Bench v-two, the harder command-line benchmark with stuff like "recover this corrupted SQLite file," the same model goes from forty-seven to fifty-nine.”Episode 003 — How to Pick the Best of Sixteen Coding Agent Rollouts