Definition
Plain language
The setup in this episode where one automated research helper is turned loose on the code that makes it work, rewriting itself over and over.
As stated in the literature
A nested self-improvement system in which an AIDE-style code-optimizing agent is applied to its own scaffolding source, with accepted rewrites gated by a held-out evaluation the inner agent cannot observe.
Also called: AIDE^2, AIDE2, AIDE-squared
Why it matters: It is a concrete, small-scale test of whether an AI system can meaningfully improve its own machinery, and the hidden evaluation is what stops it from simply gaming its own scorecard.
For example, the helper might edit the very file that tells it how many attempts to make per problem, then be re-run with that new setting to see if its scores improve.
Heard on the show
“The system built for this is called AIDE squared, one agent that optimizes code, pointed at the code that makes it an agent in the first place.”Episode 276 — An AI Agent Rewrote Its Own Scaffolding For Eight Days. Here's What Survived