Definition
Plain language
A test that measures an AI's research judgment by making it direct experiments it isn't allowed to code itself.
As stated in the literature
Benchmark separating a Researcher role (chooses experiments) from a Coder role (implements them), with monitors enforcing the boundary and scoring based on compute-matched experimental outcomes against expert human baselines.
Why it matters: It isolates research judgment from coding skill, so you can tell whether a system is genuinely good at deciding what to investigate or just good at typing out implementations.
For example, the model has to write instructions like 'try this learning rate on the smaller dataset first' and hand them to a separate coder, with a monitor rejecting any attempt to write the code itself.
Heard on the show
“Today we're discussing TasteVal, a preprint by Oliver Jaffe and Dane Sherburn at P-Zero Research.”Episode 287 — Can You Measure Research Taste If The AI Isn't Allowed To Code?