The Information Machine
Updated today·New·first covered 11 Sep 2026·5 sources

AI agent long-horizon performance collapse

The gist

Multiple benchmarks confirm AI agent success rates collapse at longer tasks

Agents that appear reliable on short benchmark tasks may fail routinely in production workflows requiring many dependent steps. Some companies report agent failures on production workloads even when benchmark scores improve, indicating a gap between controlled and deployed performance.

The full picture

Microsoft's ToolQA study, testing nine models, found agent success rates near-perfect on short runs collapse to 0-33% by 16 steps; shortening context worsened performance, ruling out context trimming as a fix. Scale AI's SWE-Bench Pro showed models scoring above 70% on standard SWE-bench drop to roughly 23% on long-horizon tasks involving multi-file refactors and cross-repository changes. METR's July 2025 time-horizon analysis placed the 50% success threshold for frontier models at approximately 50 minutes on human expert tasks. Some companies report agent failures on production workloads even when benchmark scores improve.

A Forethought analysis fit an exponential decay model to research-engineering task data, finding that success rates decline at a constant per-minute failure rate, allowing each agent to be characterized by a half-life -- the task length at which it succeeds 50% of the time, which is also the median remaining lifespan from any mid-task point. The HORIZON benchmark analyzed over 3,100 trajectories and found long-horizon failure is a structural shift in failure composition, not just a performance drop; performance gaps between frontier and smaller models collapse once agents enter this failure regime. Two distinct failure types require different interventions: environment disturbances need monitoring and recovery mechanisms, while catastrophic forgetting requires improved memory and constraint tracking. Controlling for task length, models perform worse on tasks with higher messiness scores, where messiness captures factors including resource limits, novelty, and dynamic or not-easily-resettable environments.

Anthropic reported autonomous AI agents that propose ideas, run experiments, and iterate on research problems -- specifically how to train a strong model using only a weaker model's supervision -- and found they outperform human researchers, suggesting this kind of research automation is practical.

How it developed
11 September 2026

Multiple benchmarks published September 11 documented AI agent success rates collapsing with task length, with ToolQA finding nine models fall to 0-33% by 16 steps and HORIZON finding a structural shift in failure type.

ByteDance's HarnessDev added cross-model portability as a failure mode, finding an Opus-built harness falls from 69.3 to 33.0 on SWE-Pro when a different model executes it; a Harness-of-Harness approach showed structured state-threading raises scores on extended tasks. Anthropic reported autonomous agents outperforming human researchers at training strong models from weak supervision.

9 September 2026

Microsoft and Tsinghua University paper published, showing structured run traces raise exact failure localization from 3.63% to 31.35% on GPT-4.1

7 September 2026

Microsoft ToolQA paper published: nine models collapse to 0-33% success at 16 steps from near-perfect short-run performance

5 September 2026

Harness-of-Harness (HoH) framework reported: scored 71.52 vs 58.24 for naive continuation on GameCraft-Bench after three passes

4 September 2026

ByteDance published HarnessDev, finding LLM-generated harnesses co-adapt to the creating model and lose performance when transferred to a different executor

Sources
The daily email

Want this in your inbox?

I send one email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free