The Information Machine
Following·since 23 Aug 2026·Day 6·2 sources·updated 26 Aug 2026

NVIDIA AVO and Prime Agent on ARC-AGI-3

The gist

Prime Agent Scores 95.5% on ARC-AGI-3; NVIDIA AVO Clears All 183 Levels

Two agent systems have achieved near-complete or complete scores on ARC-AGI-3 by focusing on harness design. The results indicate that how an agent manages state and memory can determine benchmark outcomes as much as the underlying model.

The full picture

Prime Agent, a self-improving agent harness paired with Claude Opus 5, scored 95.5% on ARC-AGI-3. Its technical report, titled 'Prime Agent: A Self-Improving RLM Harness,' describes a memory hierarchy spanning model weights, active context, a persistent IPython session, and a disk-backed store of histories, skills, and prompts; the agent moves state between those levels with code rather than letting the context window compact it away, reframing long-context work as an information-management problem. NVIDIA's AVO agent separately completed all 183 public ARC-AGI-3 levels. Both results point to harness architecture as a significant factor in benchmark performance on long-running tasks.

How it developed
26 August 2026

NVIDIA's AVO completed all 183 public ARC-AGI-3 levels, and Prime Agent, paired with Claude Opus 5, scored 95.5% and published an arXiv report on its architecture, both dated August 23-25.

Prime Agent's report describes a memory hierarchy where 'the model moves state between those levels with code instead of having it compacted away,' so long-context work becomes an information-management problem. Both results are framed as showing that agent scaffolding, not raw model capability, is the decisive factor.

25 August 2026

Prime Agent technical report published; system scored 95.5% on ARC-AGI-3 with Opus 5

24 August 2026

DataSpace benchmark published, harness shift from 30.98% to 46.34% accuracy reported

23 August 2026

NVIDIA's AVO agent completed all 183 public ARC-AGI-3 levels

22 August 2026

Task-CoEvolve benchmarking cost-reduction paper published

21 August 2026

HCL framework and 'harness-level forgetting' concept published

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free