All episodes
Episode 280 · Sep 27, 2026 · 12 min

Two Random Networks Teach Each Other To Predict Real Data

Cowsik, Dolev, Li et al.

LLM Pretraining
PaperDive — Episode 280: Two Random Networks Teach Each Other To Predict Real Data — cover art
paperdive.ai

Two start from random , invent their own programs, and train on nothing but the output — and the resulting model gets measurably better at predicting real text, DNA, images, and speech. The trick isn't making exercises hard; it's scoring them by whether they engage the directions the learner is already moving. We walk through what actually transfers, what the controls rule out, and the sharp line between reusable sequence and world knowledge.

Key takeaways

  • Why rewarding a generator for making exercises *hard* fails, and what the authors use instead: with the learner's own recent training
  • How 'zero data' is qualified — natural data never enters updates, but web text and DNA validation scores still guide model selection
  • The two controls that matter: adaptive scales substantially faster than a fixed program , but grammar-based still beats it on text and code
  • What the generator actually discovered by round 512 — Fibonacci-like, geometric, quadratic, and cubic sequences with -wrapping arithmetic
  • The in-context addition result: wrong-then-right progression, lower four around four examples, upper four bits around eight
  • Where the episode pushes back — the ESC-50 excludes the cost of producing it, and better DNA prediction may just mean recognizing an eight-symbol alphabet

Our reservations

The catch: skills aren't facts. The hosts push back on what the gains prove — the DNA benchmark's eight-symbol alphabet, the unmeasured split between contingent information and transferable structure. listen from 09:02

Ep. 280
Two Random Networks Teach Each Other To Predict Real Data
0:00
12 min
Paper
Self-Play Pretraining with Zero Data
Venue
arXiv:2609.30063
Year
2026
Read the paper
arxiv.org/abs/2609.30063
Also available on
Apple Podcasts Spotify

Chapters

  1. 00:00Can a tutor invent lessons from nothing?
  2. 00:46What 'zero data' does and doesn't mean
  3. 01:32Everything becomes bytes
  4. 02:10Why programs instead of sequences?
  5. 03:34The reward that isn't difficulty
  6. 04:58Two controls that keep the argument honest
  7. 05:49Does the improvement actually scale?
  8. 06:52Fibonacci out of nowhere
  9. 07:30Learning a rule with frozen weights
  10. 09:02Our reservations: the catch: skills aren't facts
  11. 09:47Does it help when real data shows up?
  12. 10:47Three takeaways and one hard boundary

References in this episode

Also available as a plain-text transcript page.

0:00Paige: Imagine a hypothetical tutor who's rewarded whenever a fails to predict the next number. The tutor could print fresh random numbers forever. The exercises would stay difficult, but there'd be no hidden rule to discover. Now make the tutor and student . Could they invent useful lessons without any examples from the world?

0:19Eric: Yes, they can, provided the tutor is rewarded by something better than difficulty. A learner trained on the outputs of self-generated programs gets better at predicting real text, image pixels, and speech samples. The interesting question is what transfers between those very different settings.

0:35Paige: This is AI Papers: A Deep Dive. We're discussing “Self-Play Pretraining with Zero Data,” a by Aditya Cowsik and colleagues, from a collaboration including researchers at Stanford and Tel Aviv University.

0:47Eric: We should pin down that title. “Zero data” doesn't mean the networks learn without examples. It means they generate their own training examples. And neither network starts as a model that already carries knowledge from human text.

1:00Paige: Both start with random , and no natural data enters their weight updates. There's an important qualification, though. The researchers check how well models predict web text and DNA, and they use those scores to choose training settings. So natural data influences which models get picked, even though it doesn't supply the training examples.

1:19Eric: That makes this a controlled experiment about what buys you. Some of it must be information about the world. But perhaps some of it is practice at recognizing structure, and that part doesn't have to come from the world.

1:33Paige: Mechanically, here means predicting the next . The model assigns probabilities to the possible continuations, then adjusts its to give the actual continuation more probability. A byte has 256 possible values. The researchers encode every evaluation domain as byte sequences, so one predictor can handle all of them.

1:52Eric: And they score prediction error in per . Lower is better, because it means the model is less surprised by what comes next. You can also think of it as a compression score. But predicting image bytes isn't recognizing objects, and predicting speech samples isn't transcribing speech. These are tests of sequence prediction.

2:11Paige: To manufacture training sequences, the team uses two with the same architecture but separate . The first one, the generator, writes short programs. A small virtual machine runs those programs, and the programs can read random input . Whatever they print becomes training data for the second network, the learner. The learner sees the outputs, not the programs.

2:33Eric: Why programs, rather than having the generator write sequences directly? The programming language can, in principle, express any computable process. That gives the search a broad space, without specifying that the output should resemble English or pictures. In practice, though, execution has strict memory and time limits, so theoretical universality doesn't mean unlimited computation.

2:55Paige: The machine is also designed so that every generated instruction string can run. Unmatched loop brackets don't cause syntax errors, and execution stops once it uses up its budget. The machine also has instructions, designed by humans, that make common operations easier to express. The comparison , which sample programs without learning, get those same shortcuts, so the shortcuts can't explain the of .

3:21Eric: But expressiveness doesn't tell the generator where to search. Most possible programs won't make useful lessons. And our hypothetical tutor's problem still applies: if you reward prediction difficulty, you can end up favoring random output. What replaces that reward?

3:35Paige: What replaces it is a question about the learner rather than about difficulty: does a candidate exercise engage the directions the learner has already been learning in? To measure that, the authors compare the exercise's with how the learner's have moved, over a growing stretch of training history. A gradient describes how changing each weight would change prediction error. So, the comparison measures how strongly that exercise connects to the path the learner is already on.

4:04Eric: There's a technical distinction here. The score uses the training algorithm's per- scaling, and it takes the absolute value of the , so either sign counts. That means it isn't literally asking whether the exercise pushes the learner forward. It's closer to asking whether the exercise is sensitive, along the directions that recent learning has established.

4:24Paige: The authors' intuition is that exercises the learner has mastered produce little signal. Unlearnable noise produces changes that don't line up with sustained progress. Useful new structure should fall between those two. That's a heuristic, not a guarantee. They chose this reward because it worked in experiments, and in small-scale tests, shuffling the rewards between programs makes worse.

4:47Eric: The generator also mixes fresh programs with mutations of promising ones, and it replays older programs. So this isn't just two networks exchanging guesses. There's machinery for exploration and for memory. How do we know that adapting the curriculum adds value?

5:02Paige: We know the adaptation adds value from two controls that keep everything except the adaptation. The first samples programs from the same language without learning which ones to favor. It shorter programs more heavily, but it never adapts to the learner. Earlier research had already explored training on random computations like this. Here, as compute grows, the adaptive system improves substantially faster than that fixed program source. So access to an expressive language isn't enough.

5:29Eric: The second control is more specialized. It generates sequences from probabilistic context-free grammars, which produce hierarchical, language-shaped structure. Training on those grammars is stronger than on text and code. But self-play is substantially better on images, music, audio, and speech. So the result is broader , not a win on every domain.

5:50Paige: Across the evaluation suite, giving more training compute produces predictable drops in next- error. The reported curves take the best tested combinations of model size, training length, and , which average several models' predictions. So they aren't simply following one model as it trains longer. All the models have fewer than 25 million , and they see about four thousand bytes of context at a time.

6:14Eric: The researchers fit power laws with a floor. That means each time you multiply compute, you get a roughly consistent fractional cut in the error that remains above that fitted floor. The regularity matters, because it shows useful doesn't appear only in one lucky run or one domain.

6:31Paige: The authors compare these improvement rates with published rates for training directly on natural data, and they're broadly comparable. But this isn't a head-to-head contest. The studies differ in scale and setup, and some of the comparison values are mathematically derived. A steep improvement curve also doesn't guarantee a good destination. Self-play can't supply missing world-specific information.

6:52Eric: The programs themselves give us a concrete view of what the search finds. By training round 512, the saved generator outputs included Fibonacci-like sequences, where each new value is the sum of the previous two, wrapping around so it fits in a . The generator also found geometric, quadratic, and cubic sequences.

7:09Paige: Meanwhile, 164 million programs from the fixed random source, produced no matches for any of those four families. That's evidence that the search finds these structures efficiently, not evidence that random search could never find them. And it doesn't establish that Fibonacci-like lessons caused the improvements on natural data. The authors name that causal question as future work.

7:30Eric: We can test something more directly than spotting attractive programs: can the learner pick up a new rule from examples in its input? That's . Its stay fixed, and the examples themselves provide the information it needs to answer the next query.

7:45Paige: With enough examples, learners approach perfect accuracy on three tasks: reversing a string, stack operations, and retrieving values from a dictionary printed in the input. That's measured on the predictions that get scored. Pretraining on the fixed program source produces little effective . Grammar is strong on dictionary retrieval, but it transfers weakly to the other tasks.

8:08Eric: The addition experiment gives a more detailed picture. The task adds two values, wrapping around at 256. Across sampled trials, as demonstrations pile up, the model first favors common output bytes. Then it tends to copy bytes from the context, which is usually wrong, and after that its predictions become less confident.

8:26Paige: After roughly four examples, it starts getting the lower four of the answer right. Around eight examples, it starts getting the upper four bits right too, and then its confidence rises. Those are patterns in its answers and probability distributions, not a report of its thoughts. But they show useful rule without any extra updates.

8:45Eric: That seems stronger than saying it learned a few recurring frequencies. It's applying relationships that were supplied in the input. Still, these are deliberately mechanical tasks. We shouldn't turn success at byte addition and stack operations, into a claim that it can learn any unfamiliar task from a prompt.

9:02Paige: And some of the natural-data improvements may come from quite simple regularities. The DNA benchmark uses only eight symbols out of those 256 possible values. Just recognizing that restricted alphabet can cut prediction error substantially, without learning any richer biological structure. That's a plausible contributor, not something the experiments isolate. Better DNA prediction doesn't automatically mean biological understanding.

9:27Eric: The authors organize these findings around two resources. One is contingent information: facts specific to a particular world or dataset. The other is transferable predictive structure, like copying and recognizing repetition. Their mathematical model separates those two, but the experiments don't measure how much of ordinary belongs to each.

9:47Paige: They also ask whether helps once real data becomes available. They take a self-play model of roughly 24 million , and use it as the starting point for normal training. On the ESC-50 audio benchmark, that reaches their convergence criterion after about 320 million training . Starting from random takes about 496 million. The narrows toward the end of training.

10:09Eric: Those are from the later training, not a saving in total compute. The comparison leaves out the cost of producing the starting point. The authors argue that one starting point can be reused across domains, which is reasonable, but the experiment doesn't settle the overall economics. What it shows is faster learning afterward.

10:28Paige: Nor does it settle whether the approach keeps working at much larger scales. Absolute prediction quality is still far below a practical model. The authors frame starting from scratch as a scientific control, not necessarily the best engineering recipe. What they've shown is narrower: generated computational structure can improve prediction beyond what it was trained on.

10:47Eric: My first takeaway is about the . Difficulty alone is a poor goal for a curriculum. This particular learning-progress signal makes adaptive program generation more useful, than continuing to sample from a fixed program distribution.

11:00Paige: My second is about . The learner gets better at predicting diverse real datasets, and it gains on tasks it never trained on. That supports the idea that some benefits of are reusable sequence skills, rather than knowledge tied to one domain.

11:16Eric: The third is the boundary: reusable skills aren't world knowledge. So can two models invent useful lessons without examples from the world? Within this controlled setup, yes. They manufacture practice that transfers, but they don't manufacture the facts that only experience with the world can supply.

11:36Paige: You can find the annotated version of this episode at paperdive.ai. It has the full transcript, with every technical term tap-to-define and related papers linked by theme. If you want every major AI paper taken apart like this, daily, that's what this channel does, so subscribe and you'll get them.

11:55Eric: The script was written by OpenAI's , and then refined by Anthropic's .5. Paige and I are AI voices from . And we're not affiliated with any of those companies. The paper is “Self-Play Pretraining with Zero Data,” by Aditya Cowsik and colleagues, posted September 24th, 2026.