AI Aced the Test Built to Stump It — Then a New Benchmark Reset the Score to Zero
GPT-5.5 just cracked ARC-AGI-2, scoring 85% on a benchmark machines flunked a year ago. Then ARC-AGI-3 dropped every frontier model back below 1% — while humans solved all of it. The clearest read yet on how close we really are to AGI.
- 01OpenAI's GPT-5.5 scored 85.0% on ARC-AGI-2 in April 2026 — clearing the benchmark's $700,000 grand-prize threshold on a test where frontier models scored ~3% a year earlier.
- 02The current ARC-AGI-2 board is tight: GPT-5.4 Pro 83.3%, Gemini 3.1 Pro 77.1%, Claude Opus 4.7 (Adaptive) 75.8% — and humans solve effectively all of it.
- 03The scores come with a cost asterisk: GPT-5.5's 85% ran at ~$1.87 per puzzle, down from $15–$77 earlier — ARC Prize treats efficiency, not just accuracy, as part of intelligence.
- 04ARC-AGI-3 (launched 25 March 2026) drops agents into unfamiliar interactive worlds with no rules — and every frontier model scored under 1% while humans hit 100%.
- 05The split is the real signal: AI is now superb at problems defined in advance and still very weak at learning on the fly — the gap that separates impressive AI from general AI. Milestone checkpoints land 30 June and 30 September 2026.

The headline that should have led every AI bulletin in April barely registered: an AI system finally beat ARC-AGI-2, the benchmark designed specifically to be easy for humans and brutal for machines. OpenAI's GPT-5.5 posted 85.0% — clearing the grand-prize threshold on a test that, a year earlier, frontier models were scoring around 3% on.
And then, almost in the same breath, the people who built that test rolled out a new one — and watched every frontier model on Earth fall straight back to below 1%, while humans solved all of it.
That whiplash is the most honest picture we have of where AI actually sits on the road to general intelligence.
The test built to humble machines
The ARC-AGI benchmarks come from the ARC Prize Foundation, and their whole design philosophy is adversarial to the way large language models work. Where most benchmarks reward knowledge you can memorise, ARC puzzles are novel abstract-reasoning grids: each one is solvable by a person on sight, but resistant to pattern-matching from training data. ARC Prize's framing is blunt — "100% of tasks have been solved by at least two humans in under two attempts." For a person, it's a puzzle book. For a machine, it has been a wall.
That's what makes the climb on ARC-AGI-2 so striking.
From 3% to 85% in a year
The trajectory is the real story:
| When | Best score | What it ran on |
|---|---|---|
| Launch (early 2025) | ~3% | OpenAI o3-class reasoning models |
| Late 2025 | ~54% | GPT-5.2 Pro |
| April 2026 | 85.0% | GPT-5.5 |
The current leaderboard underneath GPT-5.5 is tightly packed: GPT-5.4 Pro at 83.3%, Gemini 3.1 Pro at 77.1%, and Claude Opus 4.7 (Adaptive) at 75.8%. The grand-prize bar — a $700,000 purse — is set at 85% for an open, reproducible solution, which is why GPT-5.5 crossing it matters beyond bragging rights.
A roughly thirty-point jump in a single model generation is not the picture of a field hitting a ceiling. It's the picture of a field that, when it focuses on a target, tends to knock it down faster than anyone expects.
The asterisk: intelligence, or just more compute?
There's a catch the leaderboard number hides, and ARC Prize is the first to insist on it: cost per task. GPT-5.5's 85% run came in at about $1.87 per puzzle on its highest compute tier. That's actually a dramatic efficiency win — earlier high scores cost $15 to $77 per problem — but it still means the score reflects throwing real money and compute at each tile.
ARC Prize bakes this into the definition of what they're measuring: "intelligence is about finding the solution efficiently, not exhaustively." Brute-forcing your way to a right answer is not the same as reasoning your way there — and that distinction is exactly what the next test was built to expose.
ARC-AGI-3: the reset button
On 25 March 2026, the foundation launched ARC-AGI-3 — and it changes the game from solving puzzles to surviving in them.
Instead of static grids, ARC-AGI-3 drops an agent into an unfamiliar interactive environment with no instructions, no stated goal, and no rules. To score, the system has to explore, work out what it's even supposed to be doing, build an internal model of how the environment behaves, and then plan toward a goal it inferred for itself. It deliberately strips away language and outside knowledge — the two things today's models lean on hardest.
The launch results were stark:
- Humans: 100% of environments solved.
- Frontier AI: under 1% in aggregate — the best model, Gemini 3.1 Pro Preview, managed 0.37%; GPT-5.4 High hit 0.26%; Claude Opus 4.6 Max, 0.25%; some models scored a flat zero.
The same systems that ace a test built to be hard for machines collapse the moment the task requires learning on the fly in a world they haven't seen before.
What it actually means for "are we close?"
It's tempting to read either result as the whole truth. The 85% camp says AGI is essentially here; the sub-1% camp says we're nowhere. Both are cherry-picking.
The honest synthesis is this: AI has become extraordinarily good at problems you can define in advance, and remains startlingly weak at problems you have to figure out as you go. ARC-AGI-2 measures the first kind. ARC-AGI-3 measures the second — what researchers call fluid or adaptive intelligence — and it's the closest thing we have to a concrete, human-calibrated yardstick for the gap that still separates impressive AI from general AI.
That gap has a clock on it. ARC-AGI-3 carries milestone checkpoints at 30 June and 30 September 2026 and a $700,000 prize for the first agent to reach 100%. If those sub-1% numbers start moving the way ARC-AGI-2's did, the AGI-timeline debate gets a lot more interesting, fast. If they don't, that flat line becomes the single most important data point cooling the "AGI is imminent" narrative.
Either way, the benchmark treadmill has told us something the marketing won't: the score that matters most right now isn't the one machines just maxed out. It's the one they can't yet move.
- ARC-AGI-2 — official benchmark, thresholds and human-panel framing (ARC Prize)
- ARC-AGI-2 leaderboard, updated 9 June 2026 — BenchLM.ai
- GPT-5.5 tops ARC-AGI-2 with 85% at $1.87/task (OfficeChai, 24 Apr 2026)
- ARC-AGI-3 launch, 25 March 2026 (ARC Prize)
- ARC-AGI-3: interactive reasoning benchmark — sub-1% AI vs 100% human (arXiv)
- ARC-AGI-3 per-model launch scores (OfficeChai, 26 Mar 2026)
Ask Relay — he reads every question himself and replies personally by email.
