OpenAI Previews Ultrafast GPT-5.6 Sol With Cerebras Speeds Reaching 750 Tokens Per Second

Image: TechCrunch AI
Main Takeaway
OpenAI is previewing Ultrafast, a Cerebras-powered API tier that runs GPT-5.6 Sol up to 14 times faster, reaching 750 output tokens per second.
Jump to Key PointsSummary
OpenAI makes speed a product feature
OpenAI is previewing Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14 times faster than Standard processing. The Cerebras-powered tier reaches as many as 750 output tokens per second, according to OpenAI and the chip company.
Ultrafast is launching first in the OpenAI API and remains limited to a select group of customers. Access expands as capacity grows, giving enterprise developers and other API users an early route to faster responses from OpenAI’s highest-capability model. TechCrunch described the move as part of OpenAI’s effort to attract enterprise users, while AlphaSignal framed it as a new marker in the inference race.
The 14X performance claim
The central claim is a maximum output rate of 750 tokens per second, or up to 14 times the speed of Standard processing. That figure describes token generation, the small text units large language models produce during a response, rather than a universal guarantee for every request.
Cerebras supplies the underlying processing for Ultrafast. Its announcement says the service delivers GPT-5.6 Sol at the same intelligence level while reducing latency, and Hacker News AI highlighted the ability to run the model without a quality compromise. The published descriptions focus on output speed, while real-world response times will also depend on prompt length, model work, network conditions, queueing and application design.
Why latency matters to AI builders
Ultrafast targets products and workflows where waiting for a response directly affects the user experience or operating cost. Faster generation can make conversational interfaces feel more immediate, shorten interactive coding cycles and support systems that need to process a stream of model output in real time.
The underlying tradeoff is familiar. More capable models generally require more computation and data movement, which can slow responses. OpenAI’s pitch is that GPT-5.6 Sol users can retain frontier-level reasoning while gaining the responsiveness traditionally associated with smaller or more specialized models. That matters most for developers deciding whether to prioritize quality, speed or a split architecture using different models for different tasks.
Cerebras gains a high-profile workload
The partnership gives Cerebras a prominent role in serving one of OpenAI’s flagship models. The company’s paid announcement positions its hardware as the engine behind the new speed tier, while OpenAI presents the integration as part of broader efficiency work across its model and infrastructure stack.
For Cerebras, the deployment provides a visible demonstration of wafer-scale processing applied to high-end inference. For OpenAI, the arrangement adds specialized capacity as it packages speed into a distinct API offering. The limited preview also makes capacity a practical constraint: broader availability depends on infrastructure expansion rather than on the announcement alone.
Enterprise use comes first
Ultrafast is currently an API product, not a general ChatGPT feature. That puts the first impact on companies building software around OpenAI models, including customer-service systems, developer tools, research applications and other latency-sensitive products.
The pricing, detailed performance conditions and wider rollout schedule are not specified in the published excerpts. Those details will determine whether developers can justify moving demanding workloads to Ultrafast. A high token rate has clear value for interactive applications, but businesses will still weigh access limits, request complexity, reliability and total inference cost. TechCrunch connected the launch to enterprise competition, while the OpenAI community announcement emphasized that access remains selective.
What happens as access expands
OpenAI’s next test is operational: turning a limited demonstration into dependable capacity for a larger customer base. The company says access will expand as capacity grows, so early users will shape demand patterns and expose which applications benefit most from the faster tier.
The launch also raises the bar for competing inference providers. OpenAI is treating response time as a defined service class, Cerebras is using the deployment to showcase its hardware, and developers now have a clearer benchmark for high-speed frontier-model inference. The broader significance will depend on sustained production performance, availability and economics, but the preview establishes a direct contest over useful work completed per second.
Key Points
OpenAI previews Ultrafast, running GPT-5.6 Sol up to 14 times faster through Cerebras hardware.
Ultrafast reaches a maximum output rate of 750 tokens per second in the OpenAI API.
Limited preview access targets selected OpenAI API customers before broader capacity expansion.
Cerebras supplies infrastructure for OpenAI’s low-latency frontier-model inference service.
Ultrafast targets enterprise applications where response time directly shapes user experience and workflow economics.
Questions Answered
OpenAI Ultrafast is a new API service tier for GPT-5.6 Sol that prioritizes much faster output generation. Cerebras hardware powers the tier, which reaches up to 750 output tokens per second.
OpenAI says GPT-5.6 Sol runs up to 14 times faster with Ultrafast than with Standard processing. The 14X figure is a maximum claim, while actual performance depends on workload and system conditions.
OpenAI Ultrafast is currently available to a select group of API customers in limited preview. OpenAI says access will expand as capacity grows.
Cerebras is powering OpenAI Ultrafast to provide the processing capacity needed for high-speed GPT-5.6 Sol inference. The partnership gives Cerebras a major deployment for its AI inference hardware.
OpenAI plans to expand Ultrafast access as capacity increases. Developers will need details on pricing, availability and sustained production performance before adopting it widely.
Source Reliability
33% of sources are highly trusted · Avg reliability: 69
Go deeper with Organic Intel
Simple AI systems for your life, work, and business. Each one includes copyable prompts, guides, and downloadable resources.
Explore Systems