The Information Machine
Following·Day 4·first covered 28 Sep 2026·17 sources

Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%

The gist

The model delivers near-Opus 5.5 performance on several benchmarks at half the list price, reducing the cost tradeoff between quality and efficiency for enterprise AI workloads. A third model, Haiku 5.5, is expected to launch soon, completing the Claude 5.5 family.

The full picture

Anthropic launched Claude Sonnet 5.5 on September 28, 2026, the second model in its Claude 5.5 family. On Terminal-Bench 4.0, it scores 70.6% against Sonnet 5's 10.3% and Opus 5.5's 66.4%. On GDPval-AA it scores 1,844, two points below Opus 5.5's 1,846 and well above Sonnet 5's 1,449. On OSWorld 2.1 it reaches 80.1% versus Opus 5.5's 81.8%, and on Chartography it scores 61.6% against Sonnet 5's 15.6%. Vellum AI described the 395-point improvement over Sonnet 5 on knowledge-work benchmarks as the largest single-generation knowledge-work gain Anthropic has reported, though that characterization comes from a single source. Sonnet 5.5 also scores 44.7% on AutomationBench, surpassing Opus 5.5 by 2.2 points. Anthropic stated that benchmark scores capture only one facet of a model's capabilities and that Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.

Per-token pricing holds at $2 input and $10 output per million tokens, unchanged from Sonnet 5 and half of Opus 5.5. Effective per-task costs drop by up to 30% because the model uses fewer tokens and tool calls rather than a lower list price. An analysis computed that 1,000 bug-fix agent runs per month cost $420 with Sonnet 5.5 versus $1,200 with Opus 5.5, with batch API reducing Sonnet 5.5's cost to $210. CodeRabbit's code-review pipeline found Sonnet 5.5's API calls cost about 40% of Sonnet 5's for equivalent reviews, a saving of roughly 60%, which CodeRabbit attributed to Sonnet 5 repeatedly reading the same files and writing long deliberations. In agentic testing, Sonnet 5.5 fixed a multi-bug Python project in three tool calls where Sonnet 5 required twelve. A comparison on LLM-stats shows Sonnet 5.5 winning 8 of 9 shared benchmarks against GPT-5.6 Sol at roughly 2.8x lower blended per-token cost, with GPT-5.6 Sol outperforming only on DeepSWE 1.1 and offering a larger context window of 1,050,000 tokens. A Vals AI comparison shows Sonnet 5.5 leading GPT-6 Sol on 18 of 19 shared benchmarks, with the largest gap of 20.35 points on Tax Agent Bench, though error ranges overlap on four of those benchmarks. The Artificial Analysis Intelligence Index placed Sonnet 5.5 at 56 points, two behind Opus 5.5.

Sonnet 5.5 is the first Sonnet to include cyber safeguards, anti-distillation classifiers, and expanded preserved thinking. On behavioral safety audits covering approximately 1,850 scenarios, Anthropic reports it is the least likely of any Anthropic model to probe container limits. Sonnet 5.5 is the first Sonnet model to beat Pokémon Red working only from screenshots. Simon Willison noted a shared bug with Opus 5.5 where maximum thinking effort exhausts the 128,000-token budget without producing output.

Anthropicmade Sonnet 5.5 the model powering the free tier on claude.ai. Simon Willison observed that ChatGPT's free tier uses Luna 5.6. Reuters reported that enterprise customers account for about 80% of Anthropic's business and that CEO Dario Amodei earlier called on the global AI community to slow the pace of releasing new capabilities over safety concerns. Real-world deployments cited at launch include Zendesk processing support tickets 20% faster, Slack using 14% fewer output tokens, and Balyasny cutting finance task token usage from 497,000 to 121,000. Barclays disclosed on October 1 that it expects Claude Code adoption to reach 50% of its developer population by end of 2026 and a majority of software engineers by 2027, and that Claude already processes approximately 120,000 emails per day in its Global Markets business and powers a knowledge assistant used by over 16,000 employees.

How it developed
1 October 2026

Barclays disclosed Claude Code adoption targets and that Claude processes approximately 120,000 emails per day in its Global Markets business.

28 September 2026

Anthropic launched Claude Sonnet 5.5, the second model in the Claude 5.5 family, with a 70.6% score on Terminal-Bench 4.0.

Sources
12 more sources
The daily email

Want this in your inbox?

I send one email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free