Claude Sonnet 5.5 Launches: 30% Faster at No Extra Cost

Anthropic released Claude Sonnet 5.5 on September 28, the second model in the Claude 5.5 family after Opus 5.5. The company's two headline numbers: more than 30% faster than Sonnet 5, and up to 30% lower cost per task for most work. API pricing is unchanged — still $2 per million input tokens, $10 per million output tokens, and $0.2 for cached reads.

The model went live the same day on Anthropic's own platform, AWS, Google Cloud, and Microsoft Azure, with the API name claude-sonnet-5-5. The zero-data-retention option remains available across all platforms as before. Independent developer Simon Willison noted that claude.ai's free tier has already switched over to Sonnet 5.5.

The official scorecard

Here are the main results Anthropic listed on its release page:

BenchmarkSonnet 5.5
Terminal-Bench 4.070.6%
OSWorld 2.1 (partial)80.1%
Humanity's Last Exam (with tools)64.5%
CursorBench 4.055.5%
FrontierCode 1.1 (Max)46.2%
GDPval-AA v2.11844

Feedback from several early customers has centered on speed. Box says the new model is 2.4 times faster than the previous version, and Slack says it got better, faster results without changing any prompts.

"Claude Sonnet 5.5 cooks. Fast at coding and can be steered quickly in iterative workflows."

— Tyler Nishida, Every

The other side of the story, from third-party testing

Artificial Analysis gave Sonnet 5.5 an intelligence index score of 56, putting it in second place, just 2 points behind Opus 5.5's 58. In its own Terminal-Bench 4.0 test, Sonnet 5.5 scored 64%, while both Opus 5.5 and GPT-6 Astra scored 60%. Its weak spot is factual questions: 54% factual accuracy versus Opus 5.5's 66%, and its score on scientific reasoning tests is also 6 points below Opus 5.5.

The cost figures don't line up with Anthropic's own claim. Running the full test suite at the highest reasoning tier (max), Artificial Analysis found Sonnet 5.5 generates about 193,000 output tokens per task on average — roughly 60% more than the max-tier runs of both Opus 5.5 and Sonnet 5, and seven times more than GPT-6 Astra. With the same per-token price but far more tokens, the cost per task lands around $7.60, about 50% higher than Sonnet 5.

The two sides are measuring different things. Anthropic is talking about "most work" — everyday coding and writing tasks at the default tier. Artificial Analysis tested the extreme case of running the thinking budget at full tilt. Willison ran into the same pattern himself: he had the model build a small WebGL app on the max tier, which burned through 128,000 tokens and cost $1.28; switching to the xhigh tier, the same task finished in 41 seconds for about $0.06. He also noted that Sonnet 5.5 shares a quirk with Opus 5.5 — on the max tier, it tends to burn through its token budget quickly.

Where it sits among peer-tier models

Looking across the lineup, the Sonnet line's positioning is now clear: deliver roughly 80–90% of Opus's capability at a lower price. Sonnet 5 launched with the pitch that agentic tasks cost about a third of what Opus charges; the 5.5 generation adds a speed boost while closing more of the benchmark gap with Opus 5.5 — on Artificial Analysis's leaderboard, the two are now just 2 points apart.

The change to the free tier matters to more people. Willison's take is that claude.ai's free users are now getting a model considerably stronger than what ChatGPT's free tier runs on the Luna series. OpenAI's GPT-6 Luna takes the low-cost route — the latest Arena data puts its per-task cost at only about $0.05. Anthropic hasn't announced Haiku 5.5 yet; Willison expects it within the next few weeks, at which point the low-cost tier will go head-to-head with Luna.

For developers, the most immediate takeaway is: don't rush to turn on the max tier. Based on the data available now, the default or xhigh tier comes closer to the "faster and cheaper" claim Anthropic is making, while a max-tier bill can end up higher than on the previous generation. Whether it's worth it for any particular task is something Anthropic hasn't broken down by tier — the only way to know is to test it yourself.

Sources: Anthropic's official release page, Simon Willison's blog, CocoLoop, Artificial Analysis benchmark report. API pricing per Anthropic's release page; per-task cost and token usage per Artificial Analysis's intelligence index at the max tier.