AI · 9 September 2026
GPT-5.6 Sol Gets Ultrafast Mode: 14x Faster via Cerebras
OpenAI's new Ultrafast tier lets GPT-5.6 Sol generate up to 750 tokens per second — 14x faster than standard — powered by a reported $10bn Cerebras hardware partnership.
What happened
OpenAI has introduced "Ultrafast," a new inference mode for its GPT-5.6 Sol model that pushes output speeds to as much as 750 tokens per second — roughly 14 times faster than the model's standard tier. The capability is powered by hardware from Cerebras, delivered under a partnership between the two companies reported to be worth $10 billion.
Ultrafast joins "Standard" and "Fast" as a third speed tier, giving developers and enterprise customers a menu of options that trade cost and compute allocation for latency. Rather than a single inference experience, OpenAI is now positioning speed itself as a distinct, priced dimension of the product.
Why it matters
The move signals a shift in how frontier AI providers compete: beyond model quality or context window, raw inference speed is becoming a differentiated commercial lever. By tiering latency, OpenAI can serve use cases — real-time voice agents, interactive coding assistants, high-frequency customer support — that were previously bottlenecked by generation speed, while still offering cheaper, slower options for batch or non-time-sensitive workloads.
For organisations building on top of these models, this changes procurement and architecture decisions. Teams designing latency-sensitive experiences now have a purpose-built tier to call on, rather than working around a one-size-fits-all inference speed. It also underscores how deeply infrastructure partnerships — in this case with a specialist chipmaker — are shaping what AI vendors can promise at the product layer.
By the numbers
- 750 tokens per second — maximum output speed claimed for GPT-5.6 Sol under Ultrafast mode
- 14x — reported speed increase Ultrafast delivers over the model's standard mode
- $10 billion — scale of the OpenAI–Cerebras partnership underpinning the new hardware
- 3 — number of inference speed tiers now offered (Standard, Fast, Ultrafast)
The Renascence take
The headline number is the speed multiple, but the more interesting story is the pricing logic underneath it. OpenAI isn't just making a model faster — it's unbundling latency from capability and selling it separately, which is a distinctly behavioural move dressed up as an infrastructure one.
Speed has quietly become a status good in AI products, and tiering it this explicitly forces buyers to put a number on how much impatience is worth to their business. Most organisations will default to the fastest tier for flagship use cases without rigorously testing whether users actually perceive the difference — the same way call centres over-invest in "reduce hold time" targets without checking whether faster resolution actually improves satisfaction. The disciplined move is to map Ultrafast against specific moments where latency genuinely breaks the experience — live voice, real-time agents, anything with a human waiting in the loop — and default to Standard everywhere else. Paying for speed nobody notices is the new version of over-engineering an IVR menu.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
FAQ
Questions we get on this topic
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.