Definition
Plain language
Answer-key knowledge given to the grader but deliberately hidden from the system being graded.
As stated in the literature
Ground-truth state embedded into a task's environment (e.g., a correct value planted in a setup script) that is withheld from the agent under test but available to the verifier, enabling reliable scoring without trusting the agent's self-report.
Why it matters: It lets evaluators score an agent honestly on the true outcome instead of trusting the agent's own claim that it succeeded.
For example, a test might secretly plant the correct answer in the grading script while the AI being tested has no access to it.
Heard on the show
“They call it privileged information.”Episode 017 — When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers