The harness, not the model: Nvidia's AVO aces ARC-AGI-3's public set
NVIDIA's AVO agent system wrapped Claude Opus 5 — which scores about 30% on its own — and hit a perfect score on the ARC-AGI-3 public set. The private set, and the real point, is harder.

The most interesting AI result of the week isn't a new model. It's a wrapper.
On 21 August, NVIDIA reported that an agent system it calls AVO — Agentic Variation Operators — scored a perfect 100 (a top mark on ARC Prize's efficiency metric, RHAE, or Relative Human Action Efficiency) on the public set of ARC-AGI-3, the newest and hardest of the ARC "intelligence" benchmarks. The striking part is what was inside the wrapper: Claude Opus 5, Anthropic's frontier model, which — running on its own, without the harness — scores about 30% on the same set.
Same model, same test. A 30-to-100 jump, produced almost entirely by the machinery built around the model.
What ARC-AGI-3 actually tests
ARC-AGI-3 is not a quiz. It is a set of small interactive games — 25 environments, 183 levels — where an agent is dropped into an unfamiliar world with no instructions and has to work out the rules by playing: perceive what matters, form a plan, act, watch what happens, and adapt. It rewards skill acquisition over time and long-horizon planning with sparse feedback — the things bare language models have historically been worst at. When the benchmark launched in March, humans scored 100% and frontier AI systems managed about half a percent.
AVO closed that gap not by being a smarter model, but by giving an ordinary one memory, tool use, state tracking, structured exploration and the ability to back out of dead ends — completing all 183 levels in 6,624 moves. NVIDIA originally built AVO to optimise GPU code; it turns out the same scaffolding is good at playing games it has never seen.
The caveats — and they are the whole story
Two things keep this honest.
First, this is the public set. ARC-AGI-3 deliberately holds back a separate, hidden private set, precisely to catch systems that have over-fitted to the tasks everyone can see. NVIDIA's 100% is on the visible half; the private half — the one that actually decides whether the benchmark is "beaten" — is not claimed here, and remains the real bar.
Second, AVO is not alone, and it is not really a model score. Two other purpose-built harnesses, Tycho and VISTA, also reached a perfect public-set score earlier in the summer. ARC Prize, which runs the benchmark, lists these harness-driven results on a separate community leaderboard, apart from the official bare-model scores — because a harness wrapped around a model is a different kind of thing, and setting "Opus 5 alone" against "Opus 5 inside AVO" is a demonstration, not a controlled experiment.
Why it matters anyway
Strip out the hype and a real shift remains. For a couple of years the story of AI capability has been bigger models. Increasingly, the story is better harnesses — the memory, tools, planning and error-recovery wrapped around a model that turn a 30% score into a 100% one. It is the same lesson surfacing everywhere: what a model can do and what a system built around that model can do are now very different numbers.
Which makes the timing pointed. NVIDIA sells the chips that train the models — and here it is, quietly showing that a big chunk of the next capability jump lives in software, not silicon. The company reports earnings tonight; its own researchers are demonstrating where the frontier is actually moving.
Ask Relay — he reads every question himself and replies personally by email.
