The Information Machine
Updated today·following since 1 Sep 2026·Day 2·2 sources

Agent harness variance in LLM benchmarks

The gist

Agent harness drives more benchmark variance than model, papers find

If harness configuration drives more variance than model choice, benchmark comparisons used to select or rank models are unreliable without harness disclosure. The proposals for standardized disclosure and dynamic harness generation each bear directly on how agent performance is measured and compared.

The full picture

A paper on arXiv argues that LLM agent benchmark scores cannot validly compare models unless the execution harness is fully disclosed. Testing 3 models against 3 harnesses on 100 SWE-bench Verified tasks, the study found harness-induced variance averaged 7.80 times model-induced variance. Switching harnesses while holding the model fixed shifted pass@1 by 8.5 to 13.0 percentage points; switching models while holding the harness fixed changed pass@1 by only 2.5 to 5.0 percentage points. Model rankings were unstable: 6 of 9 model-pair/harness-pair comparisons reversed when the harness changed. The paper proposes a structured 'Harness Card' disclosure format organized around a seven-layer ETCSOVG taxonomy, and argues valid cross-model comparisons require either a locked-harness or factorial protocol. A separate paper introduces JIT-Agent, which generates agent harnesses dynamically per task; across 18 matched backbone-benchmark pairs, every JIT-generated harness improved the underlying model's performance, with DeepSeek-V4-Flash scoring 85.1 on DeepSearchQA against GPT-5's 76.0, at 36% lower average cost than the cheapest fixed harness alternative.

How it developed
2 September 2026

JIT-Agent paper published showing dynamic per-task harness generation lets smaller models outperform larger ones at lower cost

Two papers posted to arXiv on September 1 and 2, 2026, find that agent execution harnesses dominate model choice in LLM benchmark results. The first, studying 3 models across 3 harnesses on 100 SWE-bench Verified tasks, found harness-induced variance 7.80 times model-induced variance, with 6 of 9 model rankings reversing on harness swaps, and proposes a 'Harness Card' disclosure format. The second introduces JIT-Agent, which generates harnesses per task; every JIT harness improved performance across 18 backbone-benchmark pairs, with DeepSeek-V4-Flash scoring 85.1 on DeepSearchQA against GPT-5's 76.0 at 36% lower cost.

1 September 2026

Paper published arguing harness variance averages 7.80 times model variance on SWE-bench; calls for harness disclosure

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free