Sonnet 5.5 began powering the free tier on claude.ai
Claude Sonnet 5.5 Beats Opus 5.5 on Several Benchmarks at Half the Cost
A mid-tier model priced at half of Opus 5.5 that exceeds Opus 5.5 on several benchmarks compresses the practical cost of frontier-level performance for routine software and agentic tasks. The concrete cost gap, $420 versus $1,200 per 1,000 bug-fix runs, gives organizations a direct incentive to switch from Opus.
The full picture
Anthropic released Claude Sonnet 5.5 on September 28, 2026, and benchmark results show it exceeds its own premium Opus 5.5 on several evaluations while costing roughly half as much per task. On Terminal-Bench 4.0, Sonnet 5.5 scored 70.6%, above both Opus 5.5's 66.4% and GPT-6 Astra's 57.9%. On AutomationBench, Sonnet 5.5 scored 44.7%, surpassing Opus 5.5 by 2.2 points and GPT-6 Sol by 12.7 points. On GDPval-AA, Sonnet 5.5 lands within two Elo points of Opus 5.5 (1844 vs 1846). One analysis concluded Sonnet 5.5 achieves 98% of Opus 5.5's benchmark scores overall.
Anthropic described the model as a 30%+ speed improvement over Sonnet 5 that costs up to 30% less for most work, at unchanged per-token pricing of $2 per million input tokens and $10 per million output tokens. Per-task cost savings come from efficiency gains, specifically fewer tokens and tool calls, rather than price reductions. At 1,000 bug-fix agent runs per month, one analysis calculated Sonnet 5.5 costs $420 versus $1,200 for Opus 5.5, with batch API reducing Sonnet 5.5's cost further to $210. If a workload does not see token savings, Sonnet 5.5 costs the same as Sonnet 5.
Anthropic defended Opus 5.5, stating it remains stronger for complex, open-ended work requiring sustained judgment, and that benchmark scores capture only one facet of a model's capabilities. Sonnet 5.5 is positioned for well-scoped routine work such as software debugging, coding, document production, and UI design.
Sonnet 5.5 is the first Sonnet to include cyber safeguards, anti-distillation classifiers, and expanded preserved thinking. An adjustable effort setting lets users trade reasoning depth against cost; at low or medium effort, Sonnet 5.5 can exceed Sonnet 5's top score at about one-tenth the cost per task. The model also shares a bug with Opus 5.5 where maximum thinking effort exhausts the 128,000-token budget without producing output.
Sonnet 5.5 now powers the free tier on claude.ai. Early deployment figures cited at launch showed Zendesk processed support tickets 20% faster, Slack used 14% fewer output tokens, and Balyasny cut finance task token usage from 497,000 to 121,000. Anthropic also announced Haiku 5.5 will be available in the coming weeks.
How it developed
Sources
3 more sources
Related
- Grew out ofChinese open models' lead over US labs
- Grew out ofClaude Opus 5.5 and GPT-6 Sol release
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free