OpenAI has begun previewing a new speed tier for GPT-5.6 Sol, its most capable model. OpenAI Ultrafast runs the model up to 14× faster than Standard processing, at up to 750 output tokens per second, and launches first in the OpenAI API on Cerebras’ wafer-scale chips.

Until now, real-time speed meant trading intelligence for it: teams that needed answers in moments picked a smaller or more specialised model. Ultrafast is aimed at removing that trade.
“Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.”
OpenAI
What OpenAI Ultrafast changes
Cerebras put the mode through Humanity’s Last Exam, a 2,500-question benchmark pitched at PhD level. GPT-5.6 Sol on Ultrafast answered all of them in 11 hours and 11 minutes. Anthropic’s Claude Fable 5 needed 78 hours and 27 minutes of continuous compute to reach comparable accuracy — nearly seven times as long.
Cerebras’ figures put Ultrafast 11× faster than Claude Fable 5 and 5× faster than Claude Opus 4.8 in Fast mode, based on Artificial Analysis output speeds. On GDP-Val, its benchmark for economically valuable knowledge work, the mode delivered a 5.6× end-to-end speedup with no quality degradation. Anthropic’s own Claude Fast mode does not reach the same output rates, per TechCrunch.
Who gets access first
The OpenAI Ultrafast preview group reads like a list of buyers with speed-sensitive workloads: Jane Street, Podium, Basis and Rogo, testing the mode across coding, voice support, commerce and financial research.
“The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”
John Crepezzi, Jane Street
Ultrafast is a limited preview for a select group of customers today; access expands as capacity grows, and businesses can sign up for updates.
What it means for the UAE
The UAE is one of only three regions with an OpenAI inference residency, which keeps model processing in-country — but the GPT-5.6 family does not yet run under it, per OpenAI’s help centre. Ultrafast therefore arrives through the global API, not UAE data centres. Regional teams building on the API join the same preview list, and OpenAI’s enterprise push in the region continues; Stargate UAE, the 1GW compute hub built with G42, Oracle and SoftBank, is coming online in phases.
What is OpenAI Ultrafast?
Ultrafast is a new OpenAI API tier that runs GPT-5.6 Sol up to 14× faster than Standard processing, at up to 750 output tokens per second. It is powered by Cerebras’ wafer-scale hardware and is currently in limited preview.
How does GPT-5.6 Sol on Ultrafast compare with other AI models?
In Cerebras’ testing, GPT-5.6 Sol on Ultrafast answered all 2,500 Humanity’s Last Exam questions in 11 hours and 11 minutes, versus 78 hours 27 minutes for Anthropic’s Claude Fable 5 — comparable accuracy nearly seven times faster. Cerebras also rates it 11× faster than Fable 5 and 5× faster than Claude Opus 4.8 in Fast mode.
Is OpenAI Ultrafast available in the UAE?
Not through the UAE’s inference residency, which does not yet support the GPT-5.6 family. UAE teams building on the OpenAI API can sign up for the same limited preview, and access expands as capacity grows.


















