NVIDIA AVO and Prime Agent on ARC-AGI-3
- A controlled SWE-bench study published September 1 found harness-induced variance averaged 7.80x model-induced variance, with model rankings reversing in 6 of 9 pairings when the harness changed, and called for mandatory disclosure.
- The Yuj harness study found context management alone raised F2PF from 28% to 49% across models, while ACES found skill document quality near-zero correlated with actual runtime benefit.
- Apodex 1.1 reported gains from externalizing task state, and AutoSaddler showed unvalidated harness patching can underperform the original.
Benchmark comparisons that do not disclose or control for the harness confound model capability with harness design. The convergent findings suggest published leaderboard scores may reflect harness choices as much as model capability.