The Information Machine
Following·since 1 Sep 2026·Day 5·3 sources

Fable 5.1 Launches With Benchmark Gains but Steep Reasoning Costs

The gist

Fable 5.1 outperforms its predecessor across several evaluations, but its reasoning mode scales steeply in cost and time and can exhaust output token limits before completing long generation tasks.

The full picture

Claude Fable 5.1 is available through Perplexity Computer for Pro and Max subscribers, where Perplexity reports it ranked first in their August WANDR evaluation at a score of 0.601 and $12.76 per task, 21% higher score and 37% lower cost than Fable 5. The model offers five reasoning effort levels; independent tests by Simon Willison found that at low and medium effort it appeared to skip reasoning with costs around $0.10, while max effort produced better output at $3.30 and nearly 14 minutes of runtime. A code generation experiment tasked Fable 5.1, Fable 5, and Opus 5 with recreating Lord of the Rings landmarks in Three.js; Fable 5.1 produced the most visually detailed results but hit Anthropic's 128K output ceiling mid-file after spending 102,116 tokens on reasoning, while Fable 5 was the cheapest and fastest and the only model to complete the Bag End scene in a single call. On Code Arena: WebDev, GPT-6 Astra (Max) leads at 1,797 points, with Fable 5.1 (Max) second at 1,762 and Opus 5 (Max) third at 1,688.

How it developed
5 September 2026

Code Arena: WebDev rankings place GPT-6 Astra first at 1,797, Fable 5.1 second at 1,762

2 September 2026

Three.js Lord of the Rings experiment compares Fable 5.1, Fable 5, and Opus 5 on visual detail, cost, and token limits

1 September 2026

Simon Willison publishes evaluation of Fable 5.1 across five reasoning effort levels

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free