The Information Machine
Following·since 2 Sep 2026·Day 3·3 sources

Agent harness explains more benchmark variance than model in multiple studies

The gist

Benchmark rankings used to compare AI models may reflect harness choices as much as model capability, making standard leaderboard comparisons unreliable without harness disclosure. The findings suggest a smaller, cheaper model with the right harness can outperform a larger one, which changes how practitioners should interpret benchmark results.

The full picture

Several recent papers find that the execution harness surrounding an LLM has a larger effect on agent benchmark scores than the model itself. Zhang et al. ran a factorial experiment finding harness-induced variance exceeded model-induced variance by 7.8x, and GLM-5.1 improved from 52.5% to 65.5% solely by switching harnesses. Stanford researchers developed DuMateBench, a 200-task benchmark using tasks rebuilt from real user sessions with simulated adversity, finding that with the same model (Opus-4.8), scores ranged from 0.5821 with OpenClaw to 0.8548 with DuMate, a 27.27 percentage-point gap. JIT-Agent, which generates harnesses just-in-time for each task, allowed DeepSeek-V4-Flash to score 85.1 on DeepSearchQA against GPT-5's 76.0, and achieved 36% lower cost than fixed harness alternatives. A separate arXiv paper argues that benchmark scores for LLM agents on long-horizon tasks are not valid for cross-model comparison unless the execution harness is disclosed, and proposes a Harness Card standard organized around a seven-layer ETCSOVG taxonomy.

How it developed
3 September 2026

DuMateBench paper published showing a 27.27 percentage-point performance gap from framework choice alone using the same model

2 September 2026

JIT-Agent paper published showing smaller models can beat larger ones with task-specific harnesses, with DeepSeek-V4-Flash outscoring GPT-5 on DeepSearchQA

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free