The Information Machine
Following·Day 11·first covered 3 Sep 2026·3 sources

ARC Prize to Report Both Harness Results for GPT-6 Astra on ARC-AGI-3

The gist

The same model produces scores ranging from 54.8% to 99.9% on the same benchmark depending on evaluation scaffolding, which raises questions about what published benchmark numbers actually measure. ARC Prize's decision to label and publish both harness results gives readers a clearer basis for interpreting future scores.

The full picture

ARC Prize evaluated GPT-6 Astra on ARC-AGI-3 under two harnesses and published the results on its blog, including cost figures and a policy change. Under the Standard harness at max reasoning effort, the model scored 62.7% at a cost of $26,098. Under the Provider Adapter harness at max reasoning effort, it scored 98.6% at $17,332. At high reasoning effort, the gap is wider: 99.9% under the Provider Adapter harness versus 54.8% under the Standard harness. The Standard harness provides a minimal, provider-neutral interface where the model manages its own notes; the Provider Adapter harness preserves opaque reasoning state between requests and uses compaction. ARC Prize stated it believes a future AGI should be able to solve ARC-AGI-3 under Standard harness conditions, and announced it will report both harness results on the leaderboard going forward, with each evaluation condition clearly labeled.

How it developed
8 September 2026

ARC Prize retest with stripped-down setup reproduces 62.7% score, confirming scaffolding as cause

3 September 2026

Initial reporting on GPT-6 Astra's divergent ARC-AGI-3 scores under different harnesses published

Sources
The daily email

Want this in your inbox?

I send one email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free