Definition
Plain language
A test that drops an AI into an unfamiliar little video game with no rules or goal explained, and sees whether it can work out how to win.
As stated in the literature
Interactive successor to the ARC-AGI benchmarks: agents receive frames and a legal action set, must infer both dynamics and objective, and are scored against human action counts for level completion.
Why it matters: It tests whether a system can learn the rules of a brand-new world from scratch, which is much closer to real-world competence than answering questions about things it already read about.
For example, an AI is shown a grid with a few moving shapes and a list of buttons it can press, and nobody tells it that touching the blue square ends the level — it has to figure that out by trying things.
Heard on the show
“Today we're discussing “Kepler: Auditable World Models for ARC-AGI-3,” a public report by independent researcher Wensen Wu.”Episode 285 — What a Perfect Score Hides: Auditing an AI Agent That Scored 100