The Information Machine
Following·since 2 Sep 2026·Day 2·2 sources

Two new benchmarks show AI agents fail to improve over long horizons

The gist

Both benchmarks indicate that performance over short evaluation windows poorly predicts sustained agent quality. Neither found a model that dominates across all evaluation dimensions.

The full picture

FM-Bench and E-Commerce Bench, two recently published long-horizon agent benchmarks, both find that current AI models do not learn to improve their strategies over extended simulations and that no single model dominates. FM-Bench has 15 frontier models manage a football club across 20 simulated years with 340 to 400 decision points; on one seed, year-5 rankings correlated only 0.19 with final rankings, and DeepSeek-V4-Pro led at years 5 and 10 but finished 12th overall. E-Commerce Bench, introduced by Qwen, has agents start with ¥100,000 to run online stores for 365 simulated days handling sourcing, negotiation, pricing, and inventory; almost no model tested learns to buy cheaper or improve its strategies over the full simulated year.

How it developed
3 September 2026

Qwen introduced E-Commerce Bench, finding almost no model improves strategies over 365 simulated days of online store operation

2 September 2026

FM-Bench results published, showing year-5 rankings correlate 0.19 with final rankings across 20 simulated years of football club management

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free