Third-party benchmark firm Artificial Analysis has published updated results for Claude Fable 5.1: a composite intelligence score of 66, ranking first among the 192 models it tracks. The previous generation, Fable 5, scored 62 and placed sixth.
The same report contains another set of numbers. Running the full intelligence-index evaluation suite cost $8523.16 for Fable 5.1, versus $5455.22 for Fable 5. Converted to a per-task figure, the cost rose from $3.14 to $3.69.
Same list price, longer answers
The two generations are priced identically: $10 per million input tokens and $50 per million output tokens. The extra roughly $3,000 came entirely from output length — Fable 5 generated 83 million tokens across the full benchmark suite, while Fable 5.1 generated 140 million, an increase of about 69%. Artificial Analysis puts the median across all 192 models at 71 million tokens, meaning Fable 5.1's output volume is close to double the median.
This tracks with the broader trend in reasoning models over the past year: capability gains come from letting the model "think" through more steps, and since billing runs on tokens, the bill grows along with the length of that thinking. For ordinary users on flat monthly subscriptions, this has no direct effect; for teams billed through the API, the same workload now needs a reworked budget.
Nearly five minutes to the first token
The speed figures stand out even more. Fable 5.1's output speed is 66.4 tokens per second, ranking 84th out of 192 models, below the median of 70. Its time to first token is 296.81 seconds, versus 116.06 seconds for the previous generation — more than two and a half times longer.
That latency number is essentially unusable in an interactive setting, but it measures the full reasoning process — the model thinks for a long time before producing its first character. For an unattended background job, a five-minute wait is not a problem; waiting on autocomplete in an editor is a different matter entirely.
Set next to GPT-5.6 Sol
On the same leaderboard, OpenAI's GPT-5.6 Sol scored 61 and ranked eighth, with a full benchmark run costing $2017.29, a per-task cost of $0.95, and total output of 70 million tokens, at a list price of $4 per million input tokens and $20 per million output tokens.
Doing the rough math: Fable 5.1 scores 5 points higher on the intelligence index, at 3.9 times the per-task cost and 4.2 times the total cost of running the full suite. It's hard to say from this alone which is the better deal — the two models are suited to different kinds of work, and the time saved from getting a complex task right on the first try may well outweigh the price difference. But it does explain why more teams are starting to route tasks by tier: hard problems go to the leaderboard-topper, and the other eighty percent of routine work goes to the cheaper option.
The leaderboard answers only one question: which model is smartest. Whoever is paying the bill still has to answer a second one: how much is that one point worth.
Sources: Artificial Analysis model comparison page, CocoLoop; the intelligence index score, total benchmark cost, per-task cost, output token volume and time-to-first-token figures are all drawn from the same test run published on the firm's page.