DuMateBench paper published, showing a 27.27 percentage-point performance gap from framework choice using the same model
Agent harness drives more benchmark variance than model choice
Model rankings on widely cited benchmarks can reverse when the harness changes, making head-to-head comparisons unreliable without controlling for scaffolding. The research calls for standardized harness disclosure before benchmark scores are treated as evidence of model capability.
The full picture
Several papers find that the scaffolding and harness surrounding an LLM agent drives more performance variance on agent benchmarks than the underlying model. In a 3-model, 3-harness study on 100 SWE-bench Verified tasks, harness-induced variance averaged 7.80 times greater than model-induced variance, and model rankings reversed in 6 of 9 model-pair comparisons when the harness changed. A separate experiment varying the Yuj harness while holding models fixed raised Qwen3.6's solve rate from 28% to 49% F2PF at a 20,480-token context window; the gap nearly vanished at 262,144 tokens, indicating the harness benefit concentrates when context is the binding constraint. The DuMateBench benchmark found scores for Opus-4.8 ranging from 0.5821 to 0.8548 across frameworks, a 27.27 percentage-point spread. JIT-Agent, which generates harnesses dynamically per task, let DeepSeek-V4-Flash score 85.1 on DeepSearchQA against GPT-5's 76.0, at 36% lower average cost than fixed-harness alternatives. LoopArena isolates the controller model separately from the code-writing worker; even the top controller reached only 24.69% strict success rate across 27 tasks. Researchers have proposed a 'Harness Card' disclosure format to enable valid cross-model benchmark comparisons.
How it developed
JIT-Agent paper published, showing dynamic per-task harness generation lets smaller models outperform larger ones at lower cost
Paper reporting harness-induced variance 7.80 times greater than model-induced variance on SWE-bench Verified published, calling for harness disclosure
Yuj harness study published, showing harness changes nearly doubled solve rates on SWE-bench Verified at constrained context windows
Sources
- A strong coding model is not enough if the model managing its work does not know when to redirect, verify, or stop.
- New Stanford and other top research lab paper shows that the framework around an LLM can change agent performance dramat…
- A smaller model with the right task-specific harness can beat a stronger model.
- For long-horizon agents, this paper argues the harness can matter more than the model, so benchmark scores should not be…
- Keep the model fixed, change the harness, and coding-agent results can move a lot when context gets tight.
Want this in your inbox?
I send a short email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free