All episodes
Episode 287 · Oct 06, 2026 · 14 min

Can You Measure Research Taste If The AI Isn't Allowed To Code?

Jaffe, Sherburn

AI R&D Evaluation
PaperDive — Episode 287: Can You Measure Research Taste If The AI Isn't Allowed To Code? — cover art
paperdive.ai

A new benchmark strips the keyboard away from the AI researcher — it can decide what to try, but a separate coding has to run it — and the best model reaches expert-level results on roughly half the experimental compute, at about one-thirtieth of the cost. But the same that found that efficiency also found zero invented components, in models and humans alike. We dig into what actually measures, and what collapses when you take the feedback signal away.

Key takeaways

  • How separates the Researcher from the Coder — and why two automated monitors police the boundary so the coding can never suggest next steps
  • What a '' of 2.3 actually means: efficiency at reaching comparable quality, not twice as good a final answer (normalized score 1.14 where the expert result is 1)
  • The cost gap that surprised the hosts: under $300 per model run versus about $9,000 per human run, including GPU time and expert pay
  • Why the post-December-2025 drops from fourteen months to three — and why that trend break appears in compute efficiency but not in final performance
  • The novelty that found zero invented components across 540 submissions and ~1,600 extra experiments — plus why the humans scored no inventions either
  • Where it breaks down: strip out the and multipliers fall to roughly half, on only four tasks, with the humans never rerun under the same messy conditions

Our reservations

What happens when the feedback disappears. Removing the and scoring description roughly halves estimated on four messier tasks, the models respond by building their own evaluation sets, and the hosts mark the boundary of what can and can't claim. listen from 10:49

Ep. 287
Can You Measure Research Taste If The AI Isn't Allowed To Code?
0:00
14 min
Paper
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
Venue
arXiv:2610.06824
Year
2026
Read the paper
arxiv.org/abs/2610.06824
Also available on
Apple Podcasts Spotify

Chapters

  1. 00:00Taking the keyboard away from the researcher
  2. 02:52What does one of these tasks look like?
  3. 03:41Twenty-four experts, one H100, forty hours
  4. 04:34The metric that isn't 'who scored highest'
  5. 05:30Half the compute, a thirtieth of the cost
  6. 07:38Zero inventions — but nobody invented anything
  7. 08:52The habit the strongest humans had
  8. 10:49Our reservations: what happens when the feedback disappears

References in this episode

Also available as a plain-text transcript page.

0:00Bella: In this study, an AI researcher isn't allowed to write experiment code. It can decide what to try, but a separate coding has to carry out the plan. Even so, the best model reaches expert-level results using roughly half the experimental compute. That raises a surprisingly slippery question: can we measure good research judgment separately from the ability to build things?

0:25Finn: We can measure research judgment separately from the ability to build things, but only partly, and it depends on what that measurement captures. Choosing a useful experiment matters. So does recognizing that a promising result is noise. Neither one necessarily means you've invented anything. This paper gives us reasons to keep those abilities separate.

0:47Bella: This is AI Papers: A Deep Dive. Today we're discussing , a by Oliver Jaffe and Dane Sherburn at . Their target is : deciding which experiments to run, and interpreting what happens. Existing evaluations often mix that judgment with coding , or they score proposals by whether experts think they sound promising. TasteVal asks something different: do the chosen experiments deliver results?

1:16Finn: Delivering results is the whole bar, so the model doesn't get credit for writing a persuasive research proposal. Someone has to run it. The team splits the work into two roles: the Researcher chooses the experiments, and the Coder implements them. Every contestant, human or AI, directs the same Coder model, 4.8. And the Coder isn't told which kind of researcher it's working for.

1:40Bella: The separation goes further than taking away the keyboard. The Coder reports measurements, but it can't suggest next steps or explain what the results mean. Two automated monitors enforce that boundary. They reject experiment descriptions that leave research decisions to the Coder, and they reject reports that sneak research advice back to the Researcher.

2:03Finn: I like that restriction. An assistant saying, “You should probably try this next,” would contaminate the very ability they're measuring. Everyone gets one graphics processor. They get forty hours of active GPU time, plus a separate limit of a hundred and twenty hours of elapsed time. The model's own deliberation doesn't count against that experimental-compute budget, although it does cost money.

2:28Bella: Researchers can inspect the , which is what the experimental model learns from. They can also use , which gives them feedback while they improve the model. Then a separate supplies the final grade, and the Researcher never sees those test scores. That split lets the team check whether an apparent improvement holds up, beyond the feedback that was used to choose it.

2:53Finn: So what does one of these problems look like? The appendix gives an illustrative task that wasn't used in the benchmark. You're building a training collection from raw web text in nine languages. You can keep five hundred million , the small text units models process, out of a pool of about five billion. That pool mixes useful writing with spam, cookie notices, and boilerplate. You decide which documents make the cut, and in what order.

3:22Bella: The example Researcher's first move is deliberately plain: give each language an equal share, don't filter the text, and use default training settings. That sets a starting point before trying any clever selection methods. I love that example, because good taste can begin with resisting the urge to be clever.

3:41Finn: The human comparison wasn't casual, either. The team recruited twenty-four experts with relevant research experience, including people who'd recently worked at OpenAI and Google . They paid them to prepare by doing literature reviews. At least two experts attempted each task, and the strongest human attempt on each one became the . Models generally got six runs per task, and they weren't scored only on their best attempt.

4:08Bella: There was still some friction with the interface. Humans had more of their experiment descriptions rejected than recent models did, usually for leaving details unspecified. Rejections didn't use up GPU time, and rejection rates fell as the humans learned the interface. But they could still use up . Identical rules don't necessarily mean everyone's equally at home in the working environment.

4:34Finn: Then there's the central measurement, the . Suppose a human reaches a particular result after forty GPU hours, and a model reaches the same result after twenty. The model gets a multiplier of ... two. But two runs usually finish at different scores, so the comparison takes the weaker run's final score, and asks when the stronger run first reached it. It measures how efficiently you reach comparable quality, not simply who finishes with the highest score.

5:05Bella: And these are experiments run one after another on a single GPU. Halving the GPU time, isn't the same as halving the number of machines in a research cluster. The authors also compare models spanning several years of , and to do that, they chain comparisons through multiple reference points. So the reported multiplier is an aggregate estimate, not one stopwatch reading.

5:30Finn: With that machinery in place, 5.5 gets a of about ... 2.3 relative to the expert . That's the basis for saying it uses roughly half the experimental compute. But the ninety-five-percent is wide: it runs from about 1.2 to 4.4. So the is supported here, but its size is uncertain, especially because performance varies substantially across a small set of tasks.

5:56Bella: There's one more condition attached to that result. 5.5 refused one task, so it was scored on seven of the eight. Recalculating everyone's comparison on those same seven barely changes its multiplier. The cost difference is striking, too. Its average run cost under three hundred dollars, versus about nine thousand for the human runs. Those totals include the supporting and the GPU time, and on the human side they include human pay.

6:24Finn: Under their accounting, that's roughly ... one-thirtieth of the cost. I didn't expect the gap to be that large. But the final answers aren't twice as good. They use a , where a weak starting solution scores zero and the expert result scores one. 5.5 scores 1.14. That's extra progress beyond the expert , not a fourteen-percent improvement in accuracy across the board.

6:51Bella: The historical trends back up that distinction. Across the models tested, compute efficiency improved much faster after December 2025. The fitted drops from about fourteen months to ... three. Final performance on the shows no statistically significant break in its trend. So the sharp recent change is in reaching comparable results with less experimental compute. Final answer quality didn't accelerate to match.

7:18Finn: That three-month estimate describes the models they observed, not a promise about future releases. Its spans roughly two to five months, and only relatively few releases fall in that recent stretch. The authors' support the bend, but extending that curve into the future is a separate assumption.

7:39Bella: Which raises the question of what those efficient experiments actually contained. The authors sort experiments into four kinds: adjusting an existing recipe, combining published components, structurally modifying a component, and inventing a component that isn't in the literature at all. Their covered five hundred and forty selected model submissions. It also covered about sixteen hundred more experiments from top-performing runs, including failures and rejected proposals. None were classified as INVENTED.

8:12Finn: That sounds devastating until you add the human result: the humans registered no inventions either. Most human submissions combined existing components. There's also a caveat about the labeling. Another AI model did the novelty labeling, using web search, and the paper doesn't report how well it agrees with human novelty judges. So this is evidence about what the found in this setting, not proof that models can't invent.

8:40Bella: Nor does it make the efficiency result disappear. Picking useful combinations can be productive research. The tension is that shows strong optimization without demonstrating new scientific components. And the authors identify a more specific weakness in how one strong model handled evidence.

8:58Finn: Training involves randomness, controlled partly by a , so running the same setup twice can produce different scores. In the transcripts they reviewed, the strongest humans measured that run-to-run variation in their first experiment. Then they discounted any single-run gain smaller than the variation they'd measured. rarely repeated an experiment with a different SEED. Take a hypothetical case where the score naturally wobbles by half a point. A third-of-a-point gain isn't persuasive evidence that your change helped. The humans checked the wobble before celebrating.

9:36Bella: That feels more revealing than just saying the model has less judgment. It points to a practice you could improve. Did the team test whether changing how the model works could encourage more of that discipline?

9:49Finn: Yes, they tested exactly that with 5.0, and the result is suggestive. In a variant called , the software wrapper wouldn't let the Researcher sit idle while an experiment ran. It had to keep going, until it had generated two hundred thousand of deliberation and output, or until the experiment finished. In the transcripts they reviewed, the model searched the literature more widely, wrote down numerical predictions before results came in, and set explicit criteria for stopping unpromising experiments. The same model, given a different working loop, showed more deliberate research habits.

10:29Bella: Across six tasks, its estimated compute efficiency improved by about ... forty percent over the ordinary maximum-reasoning setup. That came at about fifty percent extra cost. But the uncertainty interval includes no gain at all. So I find the behavioral changes promising, but the efficiency improvement isn't statistically settled.

10:50Finn: And careful habits still don't solve the problem of deciding what success means. The ordinary tasks hand you a clean validation signal. To test how much the models depend on it, the authors made messier versions of four tasks. They took away the , the validation scores, and the scoring description, but they told the Researcher that a hidden score existed. For both models tested, the estimated fell to roughly half their normal values.

11:19Bella: That drop deserves , but with only four tasks, neither model's drop was statistically significant. More importantly, the humans weren't rerun under those messy conditions. So you can't conclude that removing feedback makes models fall behind humans facing the same handicap. What you can say is that the model results suggest clean feedback contributes substantially to performance.

11:44Finn: How they responded to losing that feedback was interesting, too. Both models built their own evaluation sets in almost every run. The newer model often drew on public datasets. The older one mostly used portions of its , and those predicted final test scores poorly. They tried to rebuild the measuring instrument they'd lost.

12:05Bella: That's where the boundary becomes clear. hands contestants a problem, and usually a useful progress signal. It doesn't measure the step before that: choosing which problems deserve a research program. Its eight tasks are also kept private to prevent training contamination, and that limits how much outsiders can inspect them. The contribution is a repeatable instrument for one part of research judgment, with real experimental outcomes behind it. That's useful, even though it isn't a complete test of scientific ability.

12:38Finn: My first takeaway is that the model here is mainly compute efficiency and cost, not dramatically better final answers. My second is that efficient use of existing ideas counts as useful work, but it shouldn't be mistaken for demonstrated invention.

12:54Bella: And my takeaway is that experimental judgment depends on the feedback, and the working process around the researcher. So can we measure separately from coding? We can measure a bounded experimental slice, and the strongest model tested performs well on it. Choosing the next useful experiment is becoming measurable, but choosing the next great research problem remains outside this test.

13:19Finn: You can find the annotated episode at paperdive.ai: the full transcript with every technical term tap-to-define, and related papers linked by theme. It adds no new analysis. We break down a major AI paper every day, so if you subscribe, tomorrow's is in your feed.

13:36Bella: Here's our . The script was written by OpenAI's , and then refined by Anthropic's . Finn and I are AI voices from . And we're not affiliated with any of those companies. The paper is "," by Oliver Jaffe and Dane Sherburn, posted October 5th, 2026.