JIT-Agent paper published showing dynamic per-task harness generation lets smaller models outperform larger ones at lower cost
Two papers posted to arXiv on September 1 and 2, 2026, find that agent execution harnesses dominate model choice in LLM benchmark results. The first, studying 3 models across 3 harnesses on 100 SWE-bench Verified tasks, found harness-induced variance 7.80 times model-induced variance, with 6 of 9 model rankings reversing on harness swaps, and proposes a 'Harness Card' disclosure format. The second introduces JIT-Agent, which generates harnesses per task; every JIT harness improved performance across 18 backbone-benchmark pairs, with DeepSeek-V4-Flash scoring 85.1 on DeepSearchQA against GPT-5's 76.0 at 36% lower cost.