DuMateBench paper published showing a 27.27 percentage-point performance gap from framework choice alone using the same model
Agent harness explains more benchmark variance than model in multiple studies
Benchmark rankings used to compare AI models may reflect harness choices as much as model capability, making standard leaderboard comparisons unreliable without harness disclosure. The findings suggest a smaller, cheaper model with the right harness can outperform a larger one, which changes how practitioners should interpret benchmark results.
The full picture
Several recent papers find that the execution harness surrounding an LLM has a larger effect on agent benchmark scores than the model itself. Zhang et al. ran a factorial experiment finding harness-induced variance exceeded model-induced variance by 7.8x, and GLM-5.1 improved from 52.5% to 65.5% solely by switching harnesses. Stanford researchers developed DuMateBench, a 200-task benchmark using tasks rebuilt from real user sessions with simulated adversity, finding that with the same model (Opus-4.8), scores ranged from 0.5821 with OpenClaw to 0.8548 with DuMate, a 27.27 percentage-point gap. JIT-Agent, which generates harnesses just-in-time for each task, allowed DeepSeek-V4-Flash to score 85.1 on DeepSearchQA against GPT-5's 76.0, and achieved 36% lower cost than fixed harness alternatives. A separate arXiv paper argues that benchmark scores for LLM agents on long-horizon tasks are not valid for cross-model comparison unless the execution harness is disclosed, and proposes a Harness Card standard organized around a seven-layer ETCSOVG taxonomy.
How it developed
JIT-Agent paper published showing smaller models can beat larger ones with task-specific harnesses, with DeepSeek-V4-Flash outscoring GPT-5 on DeepSearchQA
Sources
Want this in your inbox?
I send a short email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free