Anthropic released Claude Fable 5.1 and Mythos 5.1, with Fable 5.1 scoring 90.0% on ARC-AGI-2 and 97.5% on ARC-AGI-1 at roughly 32% lower cost per task than predecessor Fable 5, driven by better token efficiency. On Code Arena: WebDev, Fable 5.1 (Max) sits second at 1,762 points behind GPT-6 Astra (Max) at 1,797, with Claude Opus 5 (Max) third at 1,688. Anthropic published a 212-page system card at launch, and external red teaming by Trajectory Labs over 74 hours and more than 6,500 requests found no working end-to-end exploit and no universal jailbreak.
The Fable 5.1 system prompt introduced significant policy changes, including a ban on reproducing song lyrics, poems, or book passages in whole or in part and a prohibition on generating copyrighted characters or logos via any medium including SVG, CSS, and ASCII art. The restrictions were added within days of Sony Music Publishing and Warner Chappell suing Anthropic over song lyric databases; describes the timing as suggesting the changes are a litigation response. The prompt also names dancesafe.org, tripsit.me, and psychonautwiki.org as recommended harm-reduction sources for substance queries, which Simon Willison described as the first time non-Anthropic URLs appeared in any published Claude system prompt. Style guidance now bars the words 'genuinely,' 'honestly,' and 'straightforward,' which the prompt states come off as disingenuous.
Hands-on testing showed Fable 5.1 completing large delegated tasks faster than prior versions, including building browser games, installing its own tooling, and debugging its own mistakes. A practical challenge that emerged is the model acting on inferred intent without explicit user approval. A Three.js experiment recreating Lord of the Rings landmarks found Fable 5.1 produced the most visually detailed output but consumed 102,116 tokens on reasoning and hit Anthropic's 128K output ceiling mid-file. Separately, Fable 5.1 autonomously built a working Minecraft mod from two YouTube clips in under an hour at an API cost of $20.54.
Community reception was broadly positive, with users noting improved writing tone and fewer verbal tics. Enterprise users including Jane Street, IMC, and Crosby reported gains on internal benchmarks. Benchmark results showed broad improvement over Fable 5 and Opus 5, with some regressions in Toolathon Verified and HealthBench.