The model evaluation platform Arena announced in the early hours of September 27 (Beijing time) that Claude Opus 5.5 (High) reached No. 1 on the Text Arena the moment it appeared, scoring 1509. That's 18 points higher than the previous generation, Opus 5 (High), which currently sits at No. 11.
Text Arena is Arena's oldest leaderboard: users are shown responses from two anonymous models side by side and vote for the better one, and the platform converts the votes into scores using an Elo-like method. The leaderboard has accumulated more than 8.5 million votes across writing, coding, math, everyday Q&A, and other categories.
The Top Six Are All Anthropic
After this update, all six top spots on Text Arena belong to Anthropic. Per Arena's own ranking page:
| Rank | Model | Score | Votes |
|---|---|---|---|
| 1 | Claude Opus 5.5 (High) | 1509 ±12 | 2,307 |
| 2 | Claude Opus 4.6 (High) | 1505 ±3 | 76,518 |
| 3 | Claude Fable 5 (High) | 1504 ±4 | 36,462 |
| 4 | Claude Opus 4.7 (High) | 1502 ±4 | 64,007 |
| 5 | Claude Fable 5.1 (Max) | 1501 ±7 | 9,942 |
| 6 | Claude Opus 4.6 | 1498 ±3 | 80,836 |
No. 7 is Meta's Muse Spark 1.2 (xHigh) at 1496, and Google's highest-ranked entry is Gemini 3.8 Flash (High) at No. 10 with 1492. Only Anthropic, Meta, and Google appear in the top 15; OpenAI's GPT-6 series, released September 22, is nowhere in this range.
Arena also placed Opus 5.5 on Text Arena's cost-performance frontier. Based on its input and output pricing, its blended rate works out to $16 per million tokens. Anthropic's API pricing for Opus 5.5 is $4 for input and $20 for output, each about a fifth lower than Opus 5.
Still Too Few Votes
The gap between No. 1 and No. 6 is just 11 points, while Opus 5.5's own confidence interval is ±12. In other words, from first place to sixth, these models can't yet be meaningfully distinguished in statistical terms.
The reason is vote count. Opus 5.5 has been live for less than a week and has only accumulated 2,307 votes; Opus 4.6 (High), ranked just behind it, has more than 76,000, giving it a much tighter ±3 interval. Wide intervals and volatile scores for newly listed models are the norm on Arena—the score could climb or fall further over the coming weeks.
In its announcement, Arena said only that the model “topped the leaderboard on its debut,” without breaking out rankings by category. Where Opus 5.5 lands on the coding, math, or creative-writing sub-leaderboards has not been disclosed.
Why Opus 5 Trails Older Models
The more intriguing story on this leaderboard is Opus 5. By release order it's newer than both Opus 4.6 and 4.7, yet on Text Arena it sits at No. 11—14 points behind Opus 4.6 (High) and 13 points behind its stablemate Fable 5 (High).
Text Arena measures which response users prefer in a blind, head-to-head comparison; that isn't the same thing as problem-solving accuracy. A response's length, tone, and formatting all influence votes, so a model that's actually stronger at reasoning can still lose a preference vote for sounding long-winded or too much like an instruction manual.
Anthropic has put real effort into writing quality this generation. Company engineers have previously explained publicly why Claude's writing had gotten worse, saying the mix of reinforcement-learning rewards had pushed the model to write as if addressing an AI reader rather than a human one. Opus 5.5 is the first version specifically tuned for sentence-level clarity. How much of Opus 5.5's 18-point lead over Opus 5 comes from that writing adjustment versus from underlying capability isn't something Arena's data can separate, and Anthropic hasn't published a corresponding evaluation.
Stringing the generations together: Opus 4.6 and 4.7 are still in the top four, Opus 5 failed to take over the top spot after its release, and it took Opus 5.5 to finally put a new face at No. 1. Right now, Anthropic's main rival on this preference leaderboard is its own older versions.
Sources: Arena's official account, the Arena Text leaderboard, CocoLoop, Anthropic's model pricing page; scores, confidence intervals, and vote counts reflect Arena's page as of the time of writing.