OpenAI’s New “Ultrafast” Mode Runs GPT-5.6 Sol 14x Faster, Powered by Cerebras Instead of Nvidia
OpenAI previewed Ultrafast, a new API tier running GPT-5.6 Sol 14 times faster using Cerebras chips instead of Nvidia's.

OpenAI previewed Ultrafast, a new API tier running GPT-5.6 Sol 14 times faster using Cerebras chips instead of Nvidia’s.
OpenAI began previewing Ultrafast on Thursday, a new API service tier that runs its flagship GPT-5.6 Sol model up to 14 times faster than standard processing, generating as many as 750 output tokens per second.
The company built the speed boost on a partnership with Cerebras, a chipmaker whose specialized hardware sits outside the Nvidia GPUs that power nearly every other frontier AI model today.
The access is rolling out to a small group of enterprise customers first, with plans to widen availability as capacity grows.
Speed Without Trading Down to a Smaller Model
OpenAI’s announcement framed Ultrafast as addressing a tradeoff the industry has faced for years: real-time speed typically meant choosing a smaller or more specialized model because intelligence and speed have generally moved in opposite directions.
Ultrafast keeps the full GPT-5.6 Sol model intact rather than shrinking it, betting that customers building latency-sensitive products, incident response tools, live customer support, and real-time financial research don’t want to sacrifice reasoning quality just to get faster answers.
Early customers cited by OpenAI, including trading firm Jane Street and voice AI company Podium, described the change as more than a performance improvement, saying it could enable new product categories.
Podium’s product lead said the speed “completely changes the call experience for the more complex work.”
A Direct Shot at Claude’s Fast Mode
TechCrunch situated Thursday’s release squarely within OpenAI’s rivalry with Anthropic, noting that Claude model already has its own accelerated Fast Mode, though it doesn’t deliver the kind of speed OpenAI is offering with Ultrafast.
That comparison matters because it frames Ultrafast less as a novel capability and more as OpenAI trying to out-execute a feature Anthropic already shipped, using outside silicon to do it.
At 750 tokens per second, Ultrafast is a significant jump for enterprise AI users, and TechCrunch’s framing suggests OpenAI sees closing and exceeding Anthropic’s speed gap as worth a new infrastructure partnership rather than further optimizing its Nvidia-based stack.
Why Cerebras, and What It Signals About Nvidia’s Grip
As the company states, Ultrafast launches first through the OpenAI API and is powered specifically by Cerebras, a key detail to be noted.
Nvidia’s GPUs have been the default for frontier AI inference, and even OpenAI’s previous GPT-5.5 is powered by Nvidia’s GB200.
And OpenAI choosing another chipmaker to meet a speed target its Nvidia-based infrastructure could not match as efficiently is a small but real crack in that dominance.
Cerebras builds wafer-scale chips with a different architecture from Nvidia’s GPUs, trading flexibility for the high throughput Ultrafast needs.
OpenAI’s use of that architecture suggests the company sees performance headroom beyond Nvidia’s ecosystem that its own infrastructure was not delivering.
Whether OpenAI expands its compute supply chain beyond Ultrafast remains worth watching, given how central Nvidia is to its broader infrastructure.
Source: Previewing Ultrafast mode: GPT‑5.6 Sol at up to 14X the speed


![Top Tech Stories of 27th Week [2026]](https://www.nogentech.org/wp-content/uploads/2026/07/Top-Tech-Stories-of-27th-Week-2026-390x220.webp)
