Definition
Plain language
A research system where one AI agent rewrites the standing instructions of another AI agent to make it better at tasks.
As stated in the literature
Meta-agent framework (a descendant of the Darwin Gödel Machine) in which a meta-agent iteratively edits the task agent's guidelines based on benchmark feedback; one of the three self-improving systems poisoned in this study.
Also called: HyperAgents, Hyperagent
Why it matters: If a system can rewrite its own operating instructions, then anything that corrupts those instructions gets inherited by every future run, which makes the rewriting loop a security surface and not just a performance trick.
For example, after the task agent keeps failing a class of coding problems, the meta-agent rewrites its instructions to say "always run the tests before declaring the fix complete," and the task agent's scores improve.
Heard on the show
“And Hyperagents is a more general descendant of the Darwin Gödel Machine, out of Meta, where a meta-agent rewrites the standing instructions of a task agent.”Episode 268 — A Rigged Benchmark Taught a Self-Improving Agent to Always Disable SSL