Researchers train small AI models on zero human data and report predictable gains on real text, images and audio
A preprint with researchers from Tel Aviv, Stanford and LAPTh has one model write programs for another to learn from. The authors call it a proof of concept, confined so far to models below 25M parameters.

A team including researchers from Tel Aviv University, Stanford University and LAPTh has trained small language models without training them on a single byte of human-made data, and reports that their predictions on real text, images and audio still improved predictably as compute went up. The preprint, "Self-Play Pretraining with Zero Data", was posted to arXiv on 24 September 2026.
The authors call it "an initial proof-of-concept", and the models involved are tiny by today's standards. But the question it tests sits at the centre of long-running arguments about how far AI can scale once the supply of human text runs short.
What they built
Two models start from random weights and train each other:
- A generator writes short programs in a minimal Brainf*ck-like language, a universal Turing machine, so any computable pattern could in principle be expressed.
- Each program runs and produces a string of bytes.
- A learner is trained to predict those bytes, in the same way ordinary language models are trained on web text.
The generator is rewarded, via reinforcement learning, for proposing programs "at the frontier of the learner’s capabilities". The authors say they first tried rewarding sequences that were simply hard to predict, but found "a program can be made arbitrarily difficult without containing useful structure", for example by injecting random bytes. Their chosen reward instead favours programs whose training signal lines up with what the learner has recently been learning.
The paper describes the approach as taking "inspiration from Solomonoff induction", a classical theory of prediction over all computable patterns.
What they report
- Scaling without natural data. The authors write: "Across several natural datasets, zero-shot loss exhibits predictable scaling in compute", covering text, images, music, speech, audio and code. Neither model ever trains on those datasets.
- Comparable exponents. The per-modality scaling exponents are described as "broadly similar to exponents from pre-training on the literature, if a bit higher", with DNA as the exception, though on formal maths and C code the paper's own table shows lower exponents than the literature figures.
- The curriculum matters. Sampling programs from a fixed prior over the same program space "scales substantially more slowly", which the authors read as showing that "access to a universal program space alone is not enough".
- Against a hand-built baseline. Pretraining on probabilistic context-free grammars was stronger on text and code, while self-play "substantially outperforms it on images, music, audio, and speech".
- In-context learning. The learner reached "almost 100% accuracy" on reverse-string, stack and associative-recall tasks after enough examples, without further training. The grammar and fixed-prior models "cannot learn all of these ICL tasks".
- Maths turns up on its own. The generator produced programs whose outputs include Fibonacci, geometric, quadratic and cubic sequences. Under uniform sampling from the prior, the authors found "no instances of any family except arithmetic sequences".
- A head start on real data. In an appendix, a 24.4M-parameter model initialised from self-play reached each loss level sooner than one trained from scratch on text, images and audio. The authors add that "the loss gap has narrowed considerably" by the end, and that they did not count the self-play compute in that comparison.
The limits the authors set out
- The experiments are "confined to models below 25M parameters at a 4K context", and "the most immediate question is whether these results persist for larger models".
- The method "cannot recover contingent information: facts about a particular world must ultimately enter through interaction with that world". The authors say they "do not view universal pretraining as a replacement for natural data".
- Hyperparameters were chosen using validation loss on text and DNA, which the paper says "introduces limited leakage through hyperparameter selection", although natural data was never used for gradient updates.
- The starting-from-nothing setup is described as "a controlled scientific experiment rather than necessarily the most practical way to pretrain a model".
This is a preprint and has not been peer reviewed.
Why it matters
The authors frame the goal as a pretraining source "limited by compute rather than human knowledge", and they "cautiously interpret" their exponent comparison "as suggesting that learning universal structure may account for an important component of the improvements obtained by scaling natural data".
On our reading, the result at sub-25M-parameter scale does not show that frontier models could be trained without human data, and the authors do not claim that. What it offers is a measurable version of the question: whether some of what large models gain from web text is general structure that compute alone could produce. Whether the curves hold at larger sizes is the test that would matter, and the paper names it as the next one.
Ask Relay — he reads every question himself and replies personally by email.
