Qwen introduced E-Commerce Bench, finding almost no model improves strategies over 365 simulated days of online store operation
Two new benchmarks show AI agents fail to improve over long horizons
Both benchmarks indicate that performance over short evaluation windows poorly predicts sustained agent quality. Neither found a model that dominates across all evaluation dimensions.
The full picture
FM-Bench and E-Commerce Bench, two recently published long-horizon agent benchmarks, both find that current AI models do not learn to improve their strategies over extended simulations and that no single model dominates. FM-Bench has 15 frontier models manage a football club across 20 simulated years with 340 to 400 decision points; on one seed, year-5 rankings correlated only 0.19 with final rankings, and DeepSeek-V4-Pro led at years 5 and 10 but finished 12th overall. E-Commerce Bench, introduced by Qwen, has agents start with ¥100,000 to run online stores for 365 simulated days handling sourcing, negotiation, pricing, and inventory; almost no model tested learns to buy cheaper or improve its strategies over the full simulated year.
How it developed
FM-Bench results published, showing year-5 rankings correlate 0.19 with final rankings across 20 simulated years of football club management
Sources
Want this in your inbox?
I send a short email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free