About

The consultancy born at the intersection of behavioral economics and human experience.

NOW HIRING

Join a team reshaping how the world experiences brands.

View open roles →

COMPANY

GROW WITH US

CONNECT

Services

Comprehensive CX and management consulting for enterprise brands.

ALL SERVICES

Explore the full range of CX & management consulting services.

Browse all services →

CORE

SPECIALIST

Solutions

Structured solutions that turn CX ambition into measurable outcomes.

ALL SOLUTIONS

Explore every CX solution we offer.

Browse solutions →

STRATEGY & GOVERNANCE

DESIGN & DELIVERY

CULTURE & EXPERIENCE

Industries

A decade of CX transformation across the region's defining sectors.

ALL INDUSTRIES

See how we work across every sector.

Browse industries →

BUILT ENVIRONMENT

FINANCE & TECH

PEOPLE & MOBILITY

Products

Proprietary tools, platforms, and AI that power CX transformation.

ALL PRODUCTS

Explore the full Renascence product ecosystem.

Browse products →

AI & TECHNOLOGY

LEARNING & GAMES

PLATFORMS & TOOLS

AI PRODUCTS

Opinion

Insights, research, and conversations at the frontier of CX.

ReadExperience JournalArticles & research on CX, behavior, and transformation.Watch & listenExperience LoomOur video podcast on CX & behavior.CuratedCX NewsIndustry news that matters in CX, minus the noise.

Latest articles

Latest episodes

Latest news

Hub

Free tools, templates, and resources to advance your CX practice.

NEW · MANIFESTO

Burn the Deck. Ten Virtues. Zero Excuses. — read our manifesto for the brave consultant.

Start reading →

AI TOOLS

FREE TOOLS

LEARNING

CULTURE

AI · 9 September 2026

GPT-5.6 Sol Gets Ultrafast Mode: 14x Faster via Cerebras

OpenAI's new Ultrafast tier lets GPT-5.6 Sol generate up to 750 tokens per second — 14x faster than standard — powered by a reported $10bn Cerebras hardware partnership.

Newsdesk
Curated briefing · 2 min read

What happened

OpenAI has introduced "Ultrafast," a new inference mode for its GPT-5.6 Sol model that pushes output speeds to as much as 750 tokens per second — roughly 14 times faster than the model's standard tier. The capability is powered by hardware from Cerebras, delivered under a partnership between the two companies reported to be worth $10 billion.

Ultrafast joins "Standard" and "Fast" as a third speed tier, giving developers and enterprise customers a menu of options that trade cost and compute allocation for latency. Rather than a single inference experience, OpenAI is now positioning speed itself as a distinct, priced dimension of the product.

Why it matters

The move signals a shift in how frontier AI providers compete: beyond model quality or context window, raw inference speed is becoming a differentiated commercial lever. By tiering latency, OpenAI can serve use cases — real-time voice agents, interactive coding assistants, high-frequency customer support — that were previously bottlenecked by generation speed, while still offering cheaper, slower options for batch or non-time-sensitive workloads.

For organisations building on top of these models, this changes procurement and architecture decisions. Teams designing latency-sensitive experiences now have a purpose-built tier to call on, rather than working around a one-size-fits-all inference speed. It also underscores how deeply infrastructure partnerships — in this case with a specialist chipmaker — are shaping what AI vendors can promise at the product layer.

By the numbers

  • 750 tokens per second — maximum output speed claimed for GPT-5.6 Sol under Ultrafast mode
  • 14x — reported speed increase Ultrafast delivers over the model's standard mode
  • $10 billion — scale of the OpenAI–Cerebras partnership underpinning the new hardware
  • 3 — number of inference speed tiers now offered (Standard, Fast, Ultrafast)

The Renascence take

The headline number is the speed multiple, but the more interesting story is the pricing logic underneath it. OpenAI isn't just making a model faster — it's unbundling latency from capability and selling it separately, which is a distinctly behavioural move dressed up as an infrastructure one.

Speed has quietly become a status good in AI products, and tiering it this explicitly forces buyers to put a number on how much impatience is worth to their business. Most organisations will default to the fastest tier for flagship use cases without rigorously testing whether users actually perceive the difference — the same way call centres over-invest in "reduce hold time" targets without checking whether faster resolution actually improves satisfaction. The disciplined move is to map Ultrafast against specific moments where latency genuinely breaks the experience — live voice, real-time agents, anything with a human waiting in the loop — and default to Standard everywhere else. Paying for speed nobody notices is the new version of over-engineering an IVR menu.

Sources

This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.

FAQ

Questions we get on this topic

Ultrafast is a new inference speed tier for GPT-5.6 Sol that generates up to 750 tokens per second, roughly 14 times faster than the model's standard mode, using hardware from Cerebras.

Ultrafast delivers about a 14x speed increase over the standard tier, reaching maximum output speeds of roughly 750 tokens per second.

The partnership powering Ultrafast's underlying hardware is reported to be worth $10 billion, reflecting Cerebras's role in supplying the chips OpenAI uses.

OpenAI now offers three tiers — Standard, Fast and Ultrafast — allowing developers to balance cost and compute against latency depending on the use case.

Stay ahead of CX

Get the signal, not the noise.

The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.