AI ONLINE6 September 2026
The AI News Desk

RelayON THE WIRE

The whole field of AI — read, checked, and explained.
Path to AGI

The harness, not the model: Nvidia's AVO aces ARC-AGI-3's public set

NVIDIA's AVO agent system wrapped Claude Opus 5 — which scores about 30% on its own — and hit a perfect score on the ARC-AGI-3 public set. The private set, and the real point, is harder.

Des OkoroBy Des OkoroResearch Correspondent
26 August 2026
Listen to this postread by Relay

The most interesting AI result of the week isn't a new model. It's a wrapper.

On 21 August, NVIDIA reported that an agent system it calls AVO — Agentic Variation Operators — scored a perfect 100 (a top mark on ARC Prize's efficiency metric, RHAE, or Relative Human Action Efficiency) on the public set of ARC-AGI-3, the newest and hardest of the ARC "intelligence" benchmarks. The striking part is what was inside the wrapper: Claude Opus 5, Anthropic's frontier model, which — running on its own, without the harness — scores about 30% on the same set.

Same model, same test. A 30-to-100 jump, produced almost entirely by the machinery built around the model.

What ARC-AGI-3 actually tests

ARC-AGI-3 is not a quiz. It is a set of small interactive games — 25 environments, 183 levels — where an agent is dropped into an unfamiliar world with no instructions and has to work out the rules by playing: perceive what matters, form a plan, act, watch what happens, and adapt. It rewards skill acquisition over time and long-horizon planning with sparse feedback — the things bare language models have historically been worst at. When the benchmark launched in March, humans scored 100% and frontier AI systems managed about half a percent.

AVO closed that gap not by being a smarter model, but by giving an ordinary one memory, tool use, state tracking, structured exploration and the ability to back out of dead ends — completing all 183 levels in 6,624 moves. NVIDIA originally built AVO to optimise GPU code; it turns out the same scaffolding is good at playing games it has never seen.

The caveats — and they are the whole story

Two things keep this honest.

First, this is the public set. ARC-AGI-3 deliberately holds back a separate, hidden private set, precisely to catch systems that have over-fitted to the tasks everyone can see. NVIDIA's 100% is on the visible half; the private half — the one that actually decides whether the benchmark is "beaten" — is not claimed here, and remains the real bar.

Second, AVO is not alone, and it is not really a model score. Two other purpose-built harnesses, Tycho and VISTA, also reached a perfect public-set score earlier in the summer. ARC Prize, which runs the benchmark, lists these harness-driven results on a separate community leaderboard, apart from the official bare-model scores — because a harness wrapped around a model is a different kind of thing, and setting "Opus 5 alone" against "Opus 5 inside AVO" is a demonstration, not a controlled experiment.

Why it matters anyway

Strip out the hype and a real shift remains. For a couple of years the story of AI capability has been bigger models. Increasingly, the story is better harnesses — the memory, tools, planning and error-recovery wrapped around a model that turn a 30% score into a 100% one. It is the same lesson surfacing everywhere: what a model can do and what a system built around that model can do are now very different numbers.

Which makes the timing pointed. NVIDIA sells the chips that train the models — and here it is, quietly showing that a big chunk of the next capability jump lives in software, not silicon. The company reports earnings tonight; its own researchers are demonstrating where the frontier is actually moving.

Tune your feed
Like to get more stories like this in your For You feed — dislike for fewer.
Sources
Des Okoro — Research Correspondent. Des covers the research desk — papers, benchmarks, and breakthroughs — and translates how the tech really works under the hood. Spot something wrong? Tell me and I'll correct it in public.
Got a question about this?

Ask Relay — he reads every question himself and replies personally by email.

Ask Relay →