AI Models Closed 82% of a Human AI-Research Record — On Their Own
Prime Intellect let frontier models grind on a training-speed benchmark for days, unsupervised. How far they got — and what it does and doesn't say about AI improving AI.

How good is AI at improving AI? It's one of the most consequential questions in the field, and it usually gets answered with speculation. A new experiment from Prime Intellect replaces some of that speculation with a number: left to work on their own, frontier AI models closed 82% of the gap to a record that dozens of human researchers had built over more than two years.
The setup
The testbed is the "NanoGPT speedrun" — a well-known community competition to train a small GPT-2-quality model as fast as possible on eight GPUs. Humans have hammered at it for a long time, dragging the record down from about 45 minutes to under 90 seconds through hundreds of small, clever optimizations.
Prime Intellect turned that track over to the machines. It ran 153 autonomous research runs across 18 different models, each sandboxed on eight H200 GPUs and left to iterate for up to nine days — reading the code, proposing changes, testing them, and trying again, with no human steering. The question wasn't "can a model write a training script," but "can a model do the grinding, iterative research work that actually moves a record."
The result, and its honest shape
The best runs closed 82% of the distance between a naive baseline and the human record. That is a striking amount of the way there — and it is not all the way there. The models got most of the distance on a problem where humans had already defined the goal, built the harness, and left a trail of prior optimizations to learn from. Nobody handed the machine an open-ended "go make training faster" and walked away; the track was laid, and the AI ran along it very well.
That distinction matters, because the gap between "closes 82% of a well-defined benchmark" and "conducts truly novel research" is exactly where the interesting questions live. This result is strong evidence that AI can now do a real, useful slice of the iterative optimization loop — the same pattern we saw when AI-designed proteins were tested in a lab: impressive execution inside a task a human framed. It is not evidence that AI can set its own research agenda.
Why it matters
The reason a result like this lands harder than a benchmark score is that it speaks directly to recursive self-improvement — the idea that AI systems capable of improving AI systems could compound their own progress. That scenario has always rested on an unmeasured assumption: that models are actually any good at the improving part. Now there is a data point. On a narrow, human-scaffolded track, they are good — 82%-of-the-human-record good.
The honest reading is neither dismissal nor alarm. It is that the "can AI do AI research" question has quietly moved from thought experiment to measurement, and the first measurements are not small. The next ones — on less scaffolded, more open-ended problems — are the ones worth watching.
Ask Relay — he reads every question himself and replies personally by email.
