The Information Machine
Following·since 31 Aug 2026·Day 4·2 sources

Agent harness drives more benchmark variance than model choice

The gist

Model rankings on widely cited benchmarks can reverse when the harness changes, making head-to-head comparisons unreliable without controlling for scaffolding. The research calls for standardized harness disclosure before benchmark scores are treated as evidence of model capability.

The full picture

Several papers find that the scaffolding and harness surrounding an LLM agent drives more performance variance on agent benchmarks than the underlying model. In a 3-model, 3-harness study on 100 SWE-bench Verified tasks, harness-induced variance averaged 7.80 times greater than model-induced variance, and model rankings reversed in 6 of 9 model-pair comparisons when the harness changed. A separate experiment varying the Yuj harness while holding models fixed raised Qwen3.6's solve rate from 28% to 49% F2PF at a 20,480-token context window; the gap nearly vanished at 262,144 tokens, indicating the harness benefit concentrates when context is the binding constraint. The DuMateBench benchmark found scores for Opus-4.8 ranging from 0.5821 to 0.8548 across frameworks, a 27.27 percentage-point spread. JIT-Agent, which generates harnesses dynamically per task, let DeepSeek-V4-Flash score 85.1 on DeepSearchQA against GPT-5's 76.0, at 36% lower average cost than fixed-harness alternatives. LoopArena isolates the controller model separately from the code-writing worker; even the top controller reached only 24.69% strict success rate across 27 tasks. Researchers have proposed a 'Harness Card' disclosure format to enable valid cross-model benchmark comparisons.

How it developed
3 September 2026

DuMateBench paper published, showing a 27.27 percentage-point performance gap from framework choice using the same model

2 September 2026

JIT-Agent paper published, showing dynamic per-task harness generation lets smaller models outperform larger ones at lower cost

1 September 2026

Paper reporting harness-induced variance 7.80 times greater than model-induced variance on SWE-bench Verified published, calling for harness disclosure

31 August 2026

Yuj harness study published, showing harness changes nearly doubled solve rates on SWE-bench Verified at constrained context windows

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free