NVIDIA's AVO Took Claude Opus 5 From 30% to 100% on ARC-AGI-3 — and It Barely Touched the Model

An agent harness, not a bigger base model, tripled performance on one of the field's hardest reasoning benchmarks. The result is fueling a growing argument that the next leap in AI comes from architecture, not scale.

The most consequential AI result circulating this weekend involves no new model at all. NVIDIA's AVO agent system reportedly took Claude Opus 5 from a 30% baseline to a perfect 100% on ARC-AGI-3, according to @Josephyala, who framed it as NVIDIA demonstrating "something important about AI agents." The same base model, wrapped in different infrastructure, closed the entire gap on a benchmark designed specifically to resist brute-force pattern matching.

The distinction matters because ARC-AGI-3 is meant to test fluid, novel reasoning rather than memorized capability. A jump of that magnitude from an unchanged model suggests the ceiling being measured was never the model's raw intelligence — it was the environment the model was operating in. As @marcopapa99 put it in his August 22 briefing, "the next jump is agent architecture, not just bigger base models."

Get our free daily newsletter

Get this article free — plus the lead story every day — delivered to your inbox.

Want every article and the full archive? Upgrade anytime.

No spam. Unsubscribe anytime.