On October 2, the benchmarking platform Arena added Claude Sonnet 5.5 (Max tier) and GPT-6.1 Sol (Max tier) to its Agent Arena leaderboard. The two models landed in third and fifth place, with net improvement rates of 12.52% and 11.23%. The top two spots still belong to Anthropic's own models: Claude Fable 5.1 (14.31%) and Opus 5.5 (13.82%). Fourth place went to GPT-6 Astra (12.27%).
Agent Arena scores models differently from the text arenas people are used to. It tracks real user sessions in which a model is used as an agent — calling tools, running multi-step tasks, and ultimately getting judged on whether the user ends up satisfied. The leaderboard page shows roughly 2.158 million cumulative sessions across 51 models. The net improvement rate combines tool-call reliability, task completion, and steerability; on September 30, Arena revised its steerability scoring, expanding it from simply checking whether a model fixed something after being corrected to judging every turn in a conversation where a judgment is possible — so a model that gets something right the first time now earns more credit.
Third and Fifth, a Gap That Lives Inside the Error Bars
By rank alone, Sonnet 5.5 beats both GPT-6 Astra and GPT-6.1 Sol. Add the error margins and the picture changes: Sonnet 5.5 scores 12.52% ± 3.09%, GPT-6.1 Sol scores 11.23% ± 2.76%. The two ranges overlap heavily, and the gap to fourth-place Astra is just 0.25 percentage points. Both models are newly listed, so their session samples are still small, giving them wider error bands than Opus 5.5's ± 2.17%. Based on current data, it is hard to say with confidence which of the third-through-fifth-place models is actually stronger.
The two models each lead in different sub-metrics. Sonnet 5.5 ranks first in recovery after failed Bash commands, scoring 15.05%. GPT-6.1 Sol ranks first in tool hallucination, fabricating a nonexistent tool only 0.49% of the time, and placed second in user-confirmed task completion at 15.47%.
The Bill Tells a Different Story
The two models' API prices are identical: $2 per million input tokens, $10 per million output tokens. But in agentic tasks, the per-task cost diverges by several times. Arena's cost page shows:
| Model | Net improvement rate | Cost per task |
|---|---|---|
| Claude Fable 5.1 | 14.31% | about $4.91 |
| Claude Opus 5.5 | 13.82% | about $1.56 |
| Claude Sonnet 5.5 | 12.52% | about $2.99 |
| GPT-6 Astra | 12.27% | about $3.10 |
| GPT-6.1 Sol | 11.23% | about $0.58 |
| DeepSeek V4.1 Flash | 4.02% | about $0.11 |
Same price per token, yet per-task cost differs roughly fivefold — the gap comes down to token usage. On the Max tier, Sonnet 5.5 writes longer and "thinks" longer; third-party testing has previously found it burns noticeably more tokens per task than its predecessor. The result is that it scores lower than Anthropic's own larger Opus 5.5 while spending more, and it did not make Arena's cost-efficiency frontier (the Pareto frontier) — currently drawn by four points: Fable 5.1, Opus 5.5, GPT-6.1 Sol, and DeepSeek V4.1 Flash.
GPT-6.1 Sol is the model that redrew that frontier this time. Arena's release comparison states that its median per-task cost is 39% lower than the previous GPT-6 Sol, while scoring 1.52 percentage points higher; it is 81% cheaper than GPT-6 Astra, with a score gap within 1.04 points. GPT-6.1 Sol replaced GPT-6 Sol just over a week after launch, and this leaderboard gives OpenAI a third-party footnote for a model that is "cheaper without losing points."
These cost figures also carry methodological caveats. Arena's numbers reflect sessions users initiated on its own platform — task types lean toward coding, research, and running commands, a distribution that may not match real enterprise workflows. Costs are converted from what Arena actually paid calling the API; enterprise bills that use cache discounts or bulk discounts would come in lower. Anyone using this table for a purchasing decision should first re-test on a sample of their own tasks.
Compared With the Previous Generation
Zooming out, the Sonnet line has steadily climbed the Agent Arena rankings. The previous generation, Sonnet 5, debuted in sixth place; this time, 5.5 cracked the top three. OpenAI has moved on a different rhythm: GPT-6 Sol debuted on September 25 with a net improvement rate of 9.71%; a week later, 6.1 Sol took over, gaining about one and a half points in score while cutting cost by roughly 40%.
For developers, which model to run agentic tasks on depends on volume. For occasional, complex tasks, a high-scoring, reasonably priced model like Opus 5.5 is the easier choice. For running tasks at scale and paying by usage, lower-cost options like GPT-6.1 Sol and DeepSeek V4.1 Flash make more sense. As for Sonnet 5.5, where its cost and score would land on a lower reasoning tier is still unknown — Arena has so far only tested the Max tier.
Sources: Arena Agent Arena leaderboard, Arena cost frontier page, CocoLoop, Arena's official account, Artificial Analysis; net improvement rates, error margins, and per-task costs for each model are as shown on Arena's pages.