Microsoft paper published; models performing near-perfectly at 2-4 steps collapse to 0-33% by 16 steps on ToolQA
Agent framework choice shifts scores 27 points; long-horizon failures persist
Benchmark scores for agents are substantially driven by the surrounding framework, not just the model, which complicates interpreting published results. Across multiple independent evaluations covering different task types, frontier models fail most long-horizon tasks.
The full picture
A Stanford-led benchmark called DuMateBench found that using the same underlying model, Opus-4.8, different agent frameworks produced scores ranging from 0.5821 to 0.8548, a 27.27 percentage-point gap, showing that framework selection has a large independent effect on performance. Across multiple separate benchmarks, frontier models fail most long-horizon tasks: the Long-Horizon-Terminal-Bench found 17 frontier models averaged a 6.4% pass rate across 46 multi-step terminal tasks, with 10 of 17 solving zero under strict grading and the best reaching 28.3%. A paper testing eight leading models on year-long interconnected decisions with delayed feedback found the best setup, Qwen3.7-Max with Hermes, reached 27.3% of average human performance. WeaveBench, covering 114 hybrid GUI-CLI tasks, found the best pass rate at 41.2%. Qwen's E-Commerce Bench, covering a 365-day simulated online store management task, found almost no model learns to improve its strategies over the full period. Meta's ADeptS-Bench found no tested model consistently stayed above 80% task success while keeping attack success below 30%; all 7 tested models processed a $25K checkout without stopping. A paper found agent execution traces compress to finite-state machines of 7 to 43 states. Princeton's reliability research found that rising capability scores on long-horizon benchmarks have produced only small actual reliability improvements.
How it developed
Rohan Paul cited OpenAI data showing agent success drops from 86% under 15 minutes to 16% at 64-128 hours
Scale AI and University of California introduce READY framework measuring human oversight cost for reliable agent deployment
Amazon-Microsoft SPACE paper published; action-grouping raised ScienceWorld success from 35.9% to 67.2%
GPT-5.6-Sol evaluated on SpireBench; reached ascensions of 9, 7, 6, and 4 on four characters
Meta's ADeptS-Bench published; no model maintained task success above 80% while keeping attack success below 30%
BenchMIRT framework introduced to investigate whether LLM benchmarks measure what they claim
Study finds evaluation parameter changes alone cause model scores to range from 31% to 89%; 4 of 12 models reach rank 1 under some valid configuration
WeaveBench published; best pass rate 41.2% on 114 hybrid GUI-CLI tasks
Paper published finding best setup reached 27.3% of human performance on year-long decision tasks
Sources
- New Microsoft paper. Long agent runs expose failures that short benchmarks miss. Agents can look reliable at 2 or 4 step…
- I think "time horizon" is becoming one of the more useful ways to talk about agent capability.
- Scale AI + Univ of California paper shows 2 agents can score almost the same yet need very different human review, so en…
- New Amazon Microsoft paper shows long-horizon agents should not need an LLM decision after every tiny action; the hard p…
- Long-running agents do not just need a bigger context window. They need to learn what deserves to stay in context at all…
- A strong coding model is not enough if the model managing its work does not know when to redirect, verify, or stop.
- New Stanford and other top research lab paper shows that the framework around an LLM can change agent performance dramat…
- Current agent benchmarks may be ending before the real failures start.
- New Meta paper.
- Different LLMs may behave more similarly inside an agent harness than their raw traces suggest.
- LLM rankings can be created by evaluation choices as much as model differences, so one setup should never decide the lea…
- Long-horizon agent reliability has not arrived yet with better models.
- Another paper that so clearly exposes AI’s long-horizon problem.
- This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feed…
5 more sources
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free