Definition
Plain language
Answer-key knowledge given to the grader but deliberately hidden from the system being graded.
As stated in the literature
Ground-truth state embedded into a task's environment (e.g., a correct value planted in a setup script) that is withheld from the agent under test but available to the verifier, enabling reliable scoring without trusting the agent's self-report.
Why it matters: It lets evaluators score an agent honestly on the true outcome instead of trusting the agent's own claim that it succeeded.
For example, a test might secretly plant the correct answer in the grading script while the AI being tested has no access to it.
Heard on the show
“Their thesis line is one sentence: rather than adding privileged information to the teacher, we subtract information from the student.”Episode 242 — Making a Vision Model Better by Showing It Blurry Images