OpenAI is preparing to open its Ultrafast inference mode to more developers. On September 26, outside media spotted a not-yet-enabled speed selector hidden inside the Playground for the Responses API, offering three tiers: Standard, Fast, and Ultrafast. The rollout could land around DevDay on September 29, though OpenAI has not made an announcement, and the exact scope will depend on the official word.
Ultrafast is currently available only to a small group of invited customers. If the three-tier option goes live for everyone, speed would shift from a closed testing perk into a routine billing choice developers pick task by task.
The numbers from August's preview
Ultrafast debuted on August 13 as a limited preview, available only on GPT-5.6 Sol. OpenAI said it can output up to 750 tokens per second, up to 14 times faster than Standard processing, running on Cerebras's wafer-scale chips.
Cerebras stressed that this tier is not running a distilled or quantized, scaled-down version: the model architecture, weights, precision, context configuration, and inference settings are identical to GPT-5.6 Sol on the standard endpoint. The speed gap comes from the hardware — each wafer-scale chip carries 44GB of on-chip SRAM, so weights stay resident on the chip, cutting out the time otherwise spent shuttling data between memory and compute units.
The preview's limits were spelled out clearly too: it's API-only, unavailable in ChatGPT and Codex, and scales up only as fast as compute capacity allows.
Cerebras has published two of its own comparisons. On the GDP-Val benchmark, which measures knowledge work, it says Ultrafast delivered an end-to-end speedup of 5.6 times with no quality loss; running all 2,500 questions of Humanity's Last Exam, the Ultrafast version took 11 hours and 11 minutes, versus 78 hours and 27 minutes for Claude Fable 5. Both figures come from Cerebras's own blog under conditions it set itself, and neither has been independently reproduced yet.
"It now finishes for me before I even have the opportunity to context-switch."
That's from OpenAI researcher Jeffrey Wang — meaning the result lands before he even gets a chance to switch to something else.
The third step with Cerebras
Seen together, OpenAI's collaboration with Cerebras has now reached its third milestone with Ultrafast.
In January this year, OpenAI signed a compute deal with Cerebras to deploy up to 750 megawatts of capacity in stages through 2028. February's GPT-5.3-Codex-Spark was the first result of that deal — a purpose-built, slimmed-down coding model running on Cerebras's WSE-3, exceeding 1,000 tokens per second, and reserved for Codex inside ChatGPT Pro. In April, reports surfaced that OpenAI had added roughly $20 billion in additional purchase commitments.
Codex-Spark's approach was to build a separate small model purely for speed. Ultrafast instead runs the flagship model unchanged on faster hardware. The former trades away some capability; the latter keeps full capability and shifts the cost onto compute spend instead. Whether Cerebras's capacity can support a wider rollout will decide just how far this expansion actually goes.
How the three tiers split up
With three tiers running side by side, developers can pick latency by task. For coding assistants and live customer support — cases where someone is watching the screen waiting for a result — speed directly shapes the experience; for batch jobs and evals that run overnight, trading speed for a lower price makes more sense. Splitting the same model by response speed echoes how cloud providers charge by instance tier.
Two questions remain open. The first is price: OpenAI still hasn't published a rate for Ultrafast since the preview began, and how much more it costs than Fast is unclear. The second is coverage: GPT-6 Sol and GPT-6 Astra have already rolled out one after another, but it's not yet confirmed whether every GPT-6 model will support Ultrafast at launch.
If this does go live on DevDay, the pricing page will be the first place developers look.
Sources: OpenAI Developer Community Ultrafast preview announcement, Cerebras's official blog, CocoLoop, TestingCatalog; the 750 tokens/second figure, the 14x speedup, and the GDP-Val 5.6x speedup are all self-reported by OpenAI and Cerebras.