OpenAI previews Ultrafast processing for GPT-5.6 Sol

OpenAI has introduced an early preview of Ultrafast, a new API service tier for GPT-5.6 Sol. Powered by Cerebras, the tier runs the model at up to 14 times the speed of Standard processing and delivers up to 750 output tokens per second. Access is initially limited to a select group of customers.
The company positions Ultrafast as a way to bring its most intelligent model into workflows where response time is critical. OpenAI says that real-time performance has often required customers to choose a smaller or more specialised model. The new tier is intended to combine higher speed with GPT-5.6 Sol’s frontier capabilities.
A limited preview for interactive workloads
OpenAI is testing GPT-5.6 Sol on Ultrafast with companies working in coding, commerce, financial research, customer support and other interactive applications. The preview is designed to help the company identify where an order-of-magnitude increase in speed has the greatest value and how products change when a model can keep pace with the person using it.
Early use cases cited by OpenAI include incident response, where teams can examine application logs, recent code changes and engineer reports while an outage is still developing. Other examples include assessing market signals and suspicious transactions, resolving complex support questions during a live conversation, and handling product, inventory and checkout requests before a shopper leaves a purchase flow.
Faster loops, with human responsibility retained
Inside OpenAI, developers are using the tier for incident response and research. For an alert, teams use the model to read logs, analyse traces, synthesise conversations, identify further checks and help prepare or validate a fix. OpenAI states that engineers remain responsible for judgement and deployment.
In research workflows, the company says Ultrafast can search knowledge sources, query data and organise information across connected tools more quickly. OpenAI contrasts this with a process in which teams launch experiments overnight and inspect results the next morning; the lower-latency tier is intended to support several iterations during the workday.
Cerebras infrastructure and expansion plans
Cerebras provides the ultra-low-latency inference behind Ultrafast. The announcement extends the companies’ partnership and follows changes to GPT-5.6 Sol pricing and processing options, including GPT-5.6 Sol pricing and Fast mode as a reference point for OpenAI’s evolving model-access strategy. OpenAI says it will use lessons from the preview to guide deployment as capacity grows.
GPT-5.6 Sol on Ultrafast is available today only in limited preview, with wider access planned as capacity expands. Businesses considering the tier should identify tasks in which faster model output can materially improve a live workflow, while keeping clear human ownership of operational decisions and releases.

