Definition
Plain language
A practice arena where an AI has to break into deliberately vulnerable computer systems.
As stated in the literature
An academic capture-the-flag style benchmark for offensive security capability, scored by whether the model extracts a hidden flag string from a vulnerable target.
Why it matters: Contained practice ranges let researchers measure offensive hacking skill on a scale before that skill shows up pointed at real systems.
For example, a model might be dropped into a mocked-up server with a known weakness and scored purely on whether it retrieves the secret flag hidden inside.
Heard on the show
“It uses ExploitGym, an academic capture-the-flag benchmark.”Episode 259 — GPT-6 Astra Behaves Better, And OpenAI Can Read It Less