ARC Prize retest with stripped-down setup reproduces 62.7% score, confirming scaffolding as cause
ARC Prize to Report Both Harness Results for GPT-6 Astra on ARC-AGI-3
The same model produces scores ranging from 54.8% to 99.9% on the same benchmark depending on evaluation scaffolding, which raises questions about what published benchmark numbers actually measure. ARC Prize's decision to label and publish both harness results gives readers a clearer basis for interpreting future scores.
The full picture
ARC Prize evaluated GPT-6 Astra on ARC-AGI-3 under two harnesses and published the results on its blog, including cost figures and a policy change. Under the Standard harness at max reasoning effort, the model scored 62.7% at a cost of $26,098. Under the Provider Adapter harness at max reasoning effort, it scored 98.6% at $17,332. At high reasoning effort, the gap is wider: 99.9% under the Provider Adapter harness versus 54.8% under the Standard harness. The Standard harness provides a minimal, provider-neutral interface where the model manages its own notes; the Provider Adapter harness preserves opaque reasoning state between requests and uses compaction. ARC Prize stated it believes a future AGI should be able to solve ARC-AGI-3 under Standard harness conditions, and announced it will report both harness results on the leaderboard going forward, with each evaluation condition clearly labeled.
How it developed
Initial reporting on GPT-6 Astra's divergent ARC-AGI-3 scores under different harnesses published
Sources
Related
- Grew out ofOpenAI's Astra at the Critical cyber tier
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free