The Ultrafast inference tier has exited its invitation-only preview. On October 8, OpenAI made the service tier generally available for GPT-6.1 Sol inside the Responses API, and for the first time listed its price on the pricing page. Developers still call the model as gpt-6.1-sol and simply set service_tier to "ultrafast". The underlying model hasn't changed; what's shorter is the gap between output tokens.
The tier is open to all API users, subject to rate limits, with global processing and data residency support for both the US and the EU. Codex and ChatGPT Work have rolled it out at the same time, though access there is limited to Pro 500 users, qualifying enterprise usage-based plans, and credit-based education plans; enterprise accounts need an admin to turn it on manually.
How the five tiers are priced
Per OpenAI's pricing page, inputs up to 272K tokens count as short context, with per-tier prices as follows (US$ per million tokens):
| Tier | Input | Cached input | Cache write | Output |
|---|---|---|---|---|
| Batch / Flex | 1 | 0.05 | 1.25 | 5 |
| Standard | 2 | 0.10 | 2.50 | 10 |
| Fast | 4 | 0.20 | 5 | 20 |
| Ultrafast | 12 | 0.60 | 15 | 60 |
Beyond 272K tokens, in the long-context range, Ultrafast rises to $24 for input and $90 for output. Every one of those unit prices is six times the Standard tier's, including the two cache-related figures.
How much faster, exactly
OpenAI's documentation only gives relative multiples, not absolute speed. OpenAI's own developer account states "up to roughly 8x"; multiple outlets citing a finer breakdown put it at up to roughly 8x in Codex and up to roughly 6x through the API. VentureBeat reported a speed of around 300 tokens per second, a figure that doesn't appear in the official documentation, so the multiple should be treated as authoritative.
Ultrafast has its own separate rate limits: 1 million tokens per minute on the Build tier, 4 million on Launch, and 40 million on Grow — none of it drawn from the Standard or Fast tier's allotment.
OpenAI describes the intended use case this way:
"Built for work where speed and intelligence make a difference: debugging an outage, agents navigating apps, and live experiences where every second counts."
Running the numbers
Say an agent task reads in 200,000 tokens and outputs 20,000, ignoring caching. Rough math puts Standard at about $0.60, Fast at about $1.20, and Ultrafast at about $3.60.
Assuming the API scenario's top speed multiple of 6x, a task that used to take 6 minutes could shrink to around 1 minute — saving 5 minutes for roughly $3 more, or about $0.60 for every minute of wait saved. That's the best case; when the actual speedup falls short of the ceiling, the price per minute saved climbs higher.
For an engineer debugging a live incident, that amount barely registers. Teams running offline batch jobs will keep using Batch and Flex, both priced at half of Standard.
This tier previously went through a preview phase of its own. Ultrafast first appeared on August 13 as a limited preview attached only to GPT-5.6 Sol, when OpenAI described it as running up to 750 tokens per second and up to 14x faster than standard processing, powered by Cerebras's wafer-scale chips. By the time it reached general availability on GPT-6.1 Sol, the official multiple had become up to 6x to 8x, and this announcement made no mention of the underlying hardware.
OpenAI hasn't disclosed Ultrafast's call volume or the share of paying users adopting it. Whether it will extend to other models such as Astra also goes unmentioned in the documentation.
Sources: OpenAI developer changelog, CocoLoop, OpenAI's API pricing page (per-tier short- and long-context prices), VentureBeat, Runtime Wire; Runtime Wire verified the rate limits and rollout scope.