Cerebras enables OpenAI’s GPT-5.6 Sol ultrafast mode at 750 tokens/sec
AI chipmaker Cerebras Systems integrates its hardware with OpenAI’s latest model to deliver real-time inference at unprecedented speed.

Cerebras Systems said on Tuesday its wafer-scale AI chips now power OpenAI’s GPT-5.6 Sol in an ultrafast mode capable of generating 750 tokens per second.
The integration leverages Cerebras’s CS-3 systems, which utilize the company’s Wafer-Scale Engine 3 (WSE-3) processors, to accelerate inference for OpenAI’s latest model. The configuration is designed for real-time applications requiring low latency and high throughput, according to Cerebras.
OpenAI’s GPT-5.6 Sol is a variant of the company’s GPT-5 series, optimized for speed and efficiency in high-demand environments such as conversational AI and enterprise workflows. Cerebras did not disclose commercial terms or deployment timelines beyond the technical capability.
The announcement underscores the growing reliance on specialized hardware to meet the computational demands of advanced AI models. Competitors including Nvidia and AMD have also developed chips tailored for AI workloads, but Cerebras emphasizes its wafer-scale architecture as a differentiator in raw performance.
Cerebras Systems, a subsidiary of Cerebras AI, is privately held and has not disclosed revenue figures. The company has previously collaborated with OpenAI on research initiatives, though this marks the first public confirmation of hardware integration for production-scale inference.
Sophie covers currency markets and central bank policy across Europe, with a focus on how rate decisions ripple through FX pairs. She has been tracking the ECB's policy path since the start of the current easing cycle.
More from Sophie Laurent →

