OpenAI previews Ultrafast processing for GPT 5.6 Sol

OpenAI has introduced Ultrafast, a preview processing mode for its GPT 5.6 Sol model that the company says operates at 14 times the speed of standard processing. The mode can deliver up to 750 output tokens per second and is initially available to a small group of customers.
OpenAI says Ultrafast is intended to accelerate work on its latest and most powerful model rather than require users to switch to a smaller or more specialised system for real-time performance. Output tokens are the distinct pieces of text an LLM generates during an interaction.
A faster mode for GPT 5.6 Sol
In its announcement, OpenAI described the goal as delivering more useful work per second. The company positioned the feature as a change in the trade-off traditionally associated with real-time AI workloads, where faster responses have often meant using a smaller or more narrowly focused model.
The launch follows OpenAI's existing Sol performance options, including OpenAI Fast mode for GPT-5.6 Sol, which introduced a Fast mode for the model alongside pricing changes for GPT-5.6 Luna and Terra. Ultrafast is presented as a separate, substantially higher-speed option for GPT 5.6 Sol.
Cerebras partnership underpins the preview
Ultrafast is powered by OpenAI's partnership with chipmaker Cerebras. The preview is not yet broadly available: OpenAI says access will expand as capacity grows, without specifying a timetable or the number of customers currently included.
OpenAI identified incident response, customer service and support, financial market analysis, and e-commerce as potential corporate applications. These are workflows in which the time between a prompt and an actionable response can affect how quickly staff can process incoming work.
What businesses should evaluate
The stated 14x processing speed and ceiling of 750 output tokens per second are vendor figures, and availability remains capacity constrained during the preview. Organisations considering the mode will need to assess it within their own tasks and operating conditions as access becomes available.
For business teams, the immediate implication is to identify latency-sensitive GPT 5.6 Sol workflows, define suitable evaluation criteria, and prepare controlled tests when Ultrafast capacity reaches their accounts.

