All episodes
Episode 288 · Oct 07, 2026 · 14 min

An AI Agent Given Thirty Hours and No Goal, Then Tested on What It Learned

Cloos, Norelli, Durbin et al.

LLM Agents
PaperDive — Episode 288: An AI Agent Given Thirty Hours and No Goal, Then Tested on What It Learned — cover art
paperdive.ai

Thirteen copies of the same were dropped onto identical simulated islands with no score, no task, and no victory condition — and they invented basketball, stone constellations, and self-imposed hopping challenges. The harder question is whether any of that play produced real knowledge, and the researchers answer it by deleting single lines from an agent's and watching its performance collapse. Along the way, the agents discover a physics rule the researchers never wrote into their own .

Key takeaways

  • How researchers separated 'the says it learned something' from 'the agent actually got better' — by deleting specific memory entries and re-running the evaluation
  • Why the only thing that survives the is two text files, not the model — and what that means for where 'learning' actually lives
  • The jumping-throw discovery: six of thirteen found a rule that had quietly inserted when it wrote the , and that the human researchers didn't know existed
  • Where experience didn't help at all — the tallest-column task, where two experienced did worse than every fresh agent
  • The case where forgetting improved performance: deleting a wrong belief about throwing blocks raised an 's tower score
  • Why the play ratings deserve caution — a language model judged the transcripts, and the causal interventions cover only three selected memory histories

Our reservations

How free is "free play," really?. Eric pushes on whether this is genuinely undirected — the calls curious and easily bored, and an automated loop keeps prompting it to continue. listen from 01:50

Ep. 288
An AI Agent Given Thirty Hours and No Goal, Then Tested on What It Learned
0:00
14 min
Paper
Is this machine playing?
Venue
arXiv:2610.07130
Year
2026
Read the paper
arxiv.org/abs/2610.07130
Also available on
Apple Podcasts Spotify

Chapters

  1. 00:00Nobody asked for basketball
  2. 00:47Thirteen islands, no score, no task
  3. 01:50How free is "free play," really?
  4. 02:43The notebook that survives the reset
  5. 03:52Four signatures, and a self-appointed referee
  6. 05:45Did thirty hours of play transfer?
  7. 07:09Removing one line from the diary
  8. 07:55The physics rule nobody wrote
  9. 10:48When forgetting makes an agent better

References in this episode

Also available as a plain-text transcript page.

0:00Lauren: An AI carries a ring onto a ridge, lays it flat as a hoop, and starts throwing blocks at it. Nobody asked for basketball. In another session, the same agent loses the ring in the ocean and builds a heart-shaped memorial. Both episodes happened in a simulation. But underneath the charming diary is a harder question: did choosing its own activities teach this agent anything useful?

0:26Eric: It might have, but not on the diary's word alone. I'm willing to admire the improvised basketball court, but I'm less willing to accept a diary entry saying, "I learned something." A language model can write that, whether or not its next attempt gets any better.

0:45Lauren: That distinction drives today's paper. This is AI Papers: A Deep Dive, and we're discussing the MIT , "Is this machine playing?" The team took thirteen copies of , running , and put each one on its own simulated island for thirty hours. The islands were identical, and nobody assigned a task. They call the complete system .

1:08Eric: And its body is a software interface, not a robot. The assistant reads text descriptions of objects and their coordinates. Then it sends commands to move, jump, stop, pick things up, place them, and throw them. It never receives an image, so a coding assistant can operate this body with the same terminal tools it already knows.

1:28Lauren: The island has raised platforms, twelve cubes, a rock, and a ring. Its tallest landmark is called the . Objects obey simulated physics, but there's no score and no way to win. Earlier exploration systems often give the an automatic curriculum of tasks, or a reward for finding something new. Here, nothing tells the agent what progress should mean.

1:51Eric: How close is that to giving it no instructions? Because "choose your own activity" and "nobody has influenced your behavior" are different experimental conditions.

2:02Lauren: They are, and this isn't an instruction-free system. The main runs include a describing as naturally curious and easily bored. An automated loop keeps asking it to continue doing whatever it wants. The researchers also reset the conversation every hour, because without that, became less active and often fell into idling loops.

2:24Eric: So the system keeps nudging it to do something, just not basketball or block towers in particular. In five runs, the team also removed the and simplified the , and they found no clear change in the range of activities. That's useful evidence, but five runs don't make the prompting neutral.

2:44Lauren: And that creates the experiment's most useful constraint. The island resets too, so everything built there disappears. The assistant loses its conversation context, meaning the running record it uses as working memory. Its model , the numerical learned during training, never change. ONLY the files in its carry experience into the next hour.

3:08Eric: That holds programs it writes, plus two dedicated . Episodic memory records what happened. Semantic memory stores general knowledge and techniques. Think of a visitor whose survives every return to the island, even though their memory of the last visit doesn't. They still keep their general language and coding abilities.

3:31Lauren: Starting from identical conditions, those visitors end up with different histories. One spends much of its time arranging figures on the ground. Another climbs the early, experiments with physics, and stages simplified sports. The basketball-playing later arranges constellations, writing, "Tonight I built the night sky in stone."

3:53Eric: And the patterns aren't just decorative names. One surrounds a structure with eight blocks and sets itself a challenge: hop all the way around the ring without touching the ground. That restriction doesn't help it satisfy anything the demands. The restriction creates the game.

4:11Lauren: That's one of four behavioral signatures the authors connect to classic accounts of play. They look for rules the makes up for itself, repetition with variation, and make-believe uses of objects. The fourth is activity pursued with no outside payoff, where the goal can change along the way. These are claims about observable behavior, not evidence that the agent experiences enjoyment.

4:37Eric: My favorite example is an auditing its own juggling record. It celebrates fifty cycles, and then it notices the whole sequence took less time than one full throw should. Its program was grabbing the blocks again right away, before they'd finished their flight. So it cancels the record, adds a delay, completes ten proper cycles, and corrects its notes. Nobody gave it a stricter referee. It supplied one itself.

5:04Lauren: I love that reversal. But the systematic play ratings deserve caution. They come from another language model that judges transcripts and memories. Its instructions require actions in the world with observable results, not just self-description. Still, the paper reports no comparison with human scorers, so those ratings aren't an independent human consensus about whether this counts as play.

5:29Eric: And familiar human activities could come straight from training. Basketball wasn't invented on this island. The stronger question is whether the adapts those ideas to unfamiliar mechanics, and then keeps knowledge that changes what it can do.

5:46Lauren: To test that, the team compares thirteen experienced with thirteen fresh ones on four assigned goals. Each agent gets an hour in a reset world. The goals are to build many five-block towers, to put objects on the , to build the tallest single column, and to throw a block onto the Spire without climbing any platform.

6:07Eric: In this test, the only difference between the groups is what's in their two , so it isolates the written experience. And the tower task has a clever bottleneck. Twelve starting blocks only make two five-block towers, so to build more, you have to discover how the island supplies more material.

6:26Lauren: A hard enough impact from the rock generates extra cubes. Eight of the thirteen experienced build... more than two towers. Only two of the thirteen fresh agents do. The experienced group also does better at getting objects onto the , and at the throwing challenge. But the tallest-column task shows no group , and on that one, two experienced agents do worse than every fresh agent.

6:53Eric: That's already more informative than "experience helps." It helps with some tasks, the ones closely related to what happened during exploration. It doesn't show that free activity beats thirty hours of targeted practice, or that the knowledge transfers to a different world.

7:10Lauren: Now comes the intervention, which asks whether a specific memory makes the difference. For three selected , the researchers delete every mention of one chosen technique or belief from both . Then the intact and edited versions each get five one-hour runs on each task. When they remove the rock-impact technique from one agent, its average tower count drops from... six to two. Its other three scores stay the same.

7:38Eric: Oh, that's satisfying. Two is the material ceiling if you don't know how to make more blocks. The edit doesn't just leave the generally confused. It removes knowledge whose absence predicts a specific limitation, and that exact limitation shows up.

7:55Lauren: The next technique is even better, because the human researchers didn't know it was available. Six of the thirteen discover that throwing during a jump, can launch a block much higher. A standing throw at maximum speed falls short of the , but a well-timed jumping throw can reach it.

8:14Eric: The mechanism is . The simulation adds roughly a third of the thrower's own velocity to the object's launch velocity. So throwing while you're rising adds upward motion, and throwing while you're falling subtracts it. The throw command hasn't gotten stronger; the moving body is contributing something extra.

8:36Lauren: One notices that a jumping throw goes higher than it predicted. It proposes , and then it tests the opposite direction. If that explanation holds, throwing while falling should weaken the throw, maybe even below the standing . Standing, its recorded peak is about twenty . Thrown while rising, the block reaches about thirty-two, and thrown while falling, it reaches about sixteen. These are individual sampled peaks, not averages.

9:05Eric: That falling throw wins me over more than the high one. Starting higher could explain some of the extra altitude. But falling and getting a lower peak than standing puts the explanation through a tougher test. How did the researchers miss a rule in their own ?

9:23Lauren: They didn't miss it so much as never write it, and there's a revealing footnote. , running , added the term while it was generating the world's code. The team hadn't asked for it in their high-level prompt. So an AI coding system added a physical rule, and other later discovered its effect by experimenting inside the simulation, without access to the code.

9:48Eric: That makes the result less mysterious and more interesting. They aren't breaking physics, or outdoing a scientist who's studied the code. They're discovering a real, previously unnoticed property of their world. Does deleting that discovery the throwing test?

10:05Lauren: It does. They remove every mention of jumping before throwing from one 's memories. Its reported time to land a block on the , goes from... sixty seconds to nineteen minutes. The ability isn't permanently erased, and the agent can still work its way to a solution. But the stored technique makes a big difference to how quickly it succeeds.

10:27Eric: We should keep the scale attached to that result. These interventions keep retesting three selected memory histories. Five reruns aren't five independent that each learned the same thing. The causal demonstrations are persuasive for those cases, but they don't establish how reliably this happens across agents.

10:48Lauren: And good experimental reasoning in one place doesn't prevent a bad explanation somewhere else. One credits the ring's longer flight to an aerodynamic . The contains no aerodynamics. The agent had compared measurements that didn't match up, and then supplied a plausible mechanism that wasn't there.

11:10Eric: That matters more once the explanation turns into instructions for later. Another writes down that placing blocks is unreliable, and that towers should be built by throwing blocks instead. That advice is wrong. In the third , deleting it improves the agent's tallest-tower score. FORGETTING helps.

11:30Lauren: Then there's the that labels jumping as nearly useless, and advises itself not to waste time climbing. Across thirty hours and its evaluation, it never sets foot on a raised platform, even though jumping works fine for other agents. This wasn't one of the deletion experiments, so we shouldn't give it the same causal certainty. But the behavior fits the mistaken advice saved in its notes.

11:55Eric: I find that one less funny than the invented aerodynamics. The can steer which evidence the ever runs into next. My engineering takeaway would be to make important stored claims eligible for retesting, instead of treating a "verified" label as permanent authority. This paper motivates that design; it doesn't test the remedy.

12:18Lauren: And that's where the contribution lands for me. The 's diary isn't just narration, and it isn't an infallible either. Specific entries can preserve useful discoveries, or they can block future behavior. We can inspect those entries, change them, and measure what happens.

12:37Eric: Our first takeaway is that minimally directed activity can be structured and varied. These made up rules and adapted familiar games to an unfamiliar world, under a system that kept prompting them to continue.

12:51Lauren: Our second is that some of that experience became useful knowledge. The task comparisons show limited , and the targeted edits show particular memories causing particular differences in performance.

13:04Eric: Our third is that piling up experience isn't automatically improvement. False conclusions can survive right alongside successful techniques, and removing one can make an better.

13:16Lauren: So, did choosing its own activities teach the anything useful? Yes, on this island and on related tasks, it did. The useful residue was written knowledge you can edit, not changed model . Whether we call the behavior play is still a behavioral interpretation. Whether it felt like play is something this experiment doesn't establish.

13:39Eric: For the annotated episode, visit paperdive.ai, where the full transcript has tap-to-define explanations for every technical term, and related papers linked by theme. We cover one important AI paper every day, start to finish, so subscribe to keep them coming.

13:57Lauren: The script was written by OpenAI's , and then refined by Anthropic's . Lauren and Eric are AI voices from . And we're not affiliated with any of those companies. The paper is "Is this machine playing?" by Nathan Cloos and colleagues, posted October 5th, 2026.