Astra, Sol Subscription Speed Jumps to 50 Tokens/Second

OpenAI has bumped the default output speed for GPT-6 Astra and GPT-6.1 Sol across its subscription products. Codex lead Thibault Sottiaux announced on social media on October 5 that the default speed for both models has increased by roughly 50% overall, with output rising from about 30 tokens per second to about 50. The change covers all of OpenAI's own products as well as every partner app connected through "Sign in with ChatGPT."

Users don't need to change any settings. Sottiaux said the update would roll out to everyone within two hours.

An "improvement" delivered the very next day

The timing of this speed boost is no coincidence. On October 4, Sottiaux set a public 28-day deadline for the Codex team: every day, the team must either ship something that noticeably improves the experience for most Codex and ChatGPT Work users, or give everyone a full usage reset.

A speed increase fits neatly within that "most users can feel it" bar. It isn't tied to any particular use case — coding, research, and agent tasks all benefit — and it requires no one to learn a new feature, which makes it an easy win in a commitment that demands something new every single day.

The reach also goes well beyond ChatGPT itself. The partners Sottiaux named include the open-source coding tool OpenCode, along with Pi, Amp, and Cognition's Devin. These products let users sign in directly with their ChatGPT subscription, drawing on subscription quota rather than API billing — so the subscription-side speed boost carries straight through to them.

30 to 50 adds up to more than half

The official headline says "about 50%," but by the two numbers Sottiaux actually gave — 30 tokens per second rising to 50 — the increase works out to roughly 67%. Sottiaux didn't explain the gap between the two figures. One possibility is that 50% refers to the end-to-end experience, including things like the wait for the first token, while the generation phase alone is about two-thirds faster.

Translated into time users would actually notice, here's a rough estimate:

Output lengthAt 30 tokens/secAt 50 tokens/sec
A 2,000-token block of codeAbout 67 secondsAbout 40 seconds
A 20,000-token agent task outputAbout 11 minutesAbout 6.7 minutes

That only accounts for the time the model spends generating text. Agent tasks also involve tool calls, running tests, and waiting on the network — none of which get faster just because output speed did — so the actual time saved on finishing a task will be lower than what the table suggests.

Tokenizer and token usage

Sottiaux also noted that both models now use an optimized tokenizer, which reduces the number of tokens needed to complete the same task. OpenAI didn't provide a specific figure for how much lower.

This point may matter more to subscribers than the speed boost itself. Codex and ChatGPT Work quotas are calculated based on token consumption, so if the same task uses fewer tokens, the same quota goes further. Conversely, speed alone doesn't change quota at all: faster output just means users hit the cap on their five-hour window sooner. Some outlets specifically flagged this in their coverage.

When GPT-6.1 Sol launched at DevDay on September 29, third-party testing found it used 10% to 30% more output tokens than the previous generation to complete the same tasks. How much the new tokenizer offsets that gap won't be clear until independent testing is run again.

No change yet on the API side

The two models have a clear division of labor. GPT-6 Astra is the flagship, priced on the API at the standard rate of $10 per million input tokens and $50 per million output tokens. GPT-6.1 Sol is aimed at coding agents and computer-use tasks, priced at $2 per million input tokens and $10 per million output tokens — one-fifth of Astra's rate on both counts. This subscription-side speed boost applies to both models at once.

This adjustment only touches the default speed in subscription products. API users pay per token, and Sottiaux's remarks didn't address whether speed changes there too; OpenAI's API documentation hasn't been updated on the matter either.

Sources: Thibault Sottiaux's public social media statement, CocoLoop, RuntimeWire, Crypto Briefing. Output rates are based on the official tokens-per-second figures given; increases and time estimates are derived from that baseline.