Glossary · Term

privileged information

← all terms

Definition

Plain language

Answer-key knowledge given to the grader but deliberately hidden from the system being graded.

As stated in the literature

Ground-truth state embedded into a task's environment (e.g., a correct value planted in a setup script) that is withheld from the agent under test but available to the verifier, enabling reliable scoring without trusting the agent's self-report.

Why it matters: It lets evaluators score an agent honestly on the true outcome instead of trusting the agent's own claim that it succeeded.

For example, a test might secretly plant the correct answer in the grading script while the AI being tested has no access to it.

Heard on the show

“Their thesis line is one sentence: rather than adding privileged information to the teacher, we subtract information from the student.”
Episode 242 — Making a Vision Model Better by Showing It Blurry Images

Mentioned in 4 episodes

  1. 242
    Making a Vision Model Better by Showing It Blurry Images
  2. 155
    Why a Flawless Demo Makes a Worse Computer-Using Agent, And the Fix
  3. 099
    How an Open-Book Trick Teaches a Model to Catch Its Own Mistakes
  4. 017
    When the Agent Grades Its Own Homework: A Brutal New Benchmark for AI Workers

Related terms