Definition
Plain language
A method that rewrites an AI agent's surrounding code to improve it, but checks each change against a set of known-correct answers.
As stated in the literature
A harness-optimization method that edits full harness code including tools, scored against a labeled validation set; used as the label-hungry head-to-head baseline against the label-free RHO.
Why it matters: It can sharply improve an agent's tooling, but only when you have a labeled set of right answers to check changes against.
For example, it might rewrite an agent's set of tools and then test each new version against a list of problems whose correct answers are already known.
Heard on the show
“Think Darwin Gödel Machine, Meta-Harness, the AI Scientist line of work.”Episode 088 — Two Levers for Self-Improving AI: When Rewriting Code Isn't Enough