OpenAI Previews Ultrafast Mode for GPT-5.6 Sol at Up to 14x Speed

OpenAI introduces Ultrafast, a preview API service tier that runs GPT-5.6 Sol up to 14x faster and delivers up to output tokens per second.

Aug 15, 2026 - 05:56
 1
OpenAI Previews Ultrafast Mode for GPT-5.6 Sol at Up to 14x Speed
Image Credit: Chatgpt

Verified against OpenAI’s official announcement and Cerebras. Ultrafast is an “OpenAI API service tier”, initially available to select customers, that runs GPT-5.6 Sol up to 14x faster than Standard processing and reaches up to 750 output tokens per second. 

OpenAI is previewing a new Ultrafast service tier that dramatically increases the speed of GPT-5.6 Sol, targeting developers and businesses that need frontier AI performance with much lower response times.

Ultrafast runs GPT-5.6 Sol at up to 14 times the speed of OpenAI’s Standard processing tier and can generate as many as 750 output tokens per second. The new option is launching first through the OpenAI API rather than as a general ChatGPT mode.

OpenAI said faster inference has traditionally required users to choose smaller or more specialised models. Ultrafast is designed to offer an alternative by retaining the capabilities of GPT-5.6 Sol while substantially increasing the speed at which the model produces results.

OpenAI targets workflows where speed matters

The company sees the service as useful for applications where delays can directly affect the user experience or business decisions. OpenAI highlighted incident response, customer service, financial analysis, and e-commerce as workloads that could benefit from faster model output.

The technology behind Ultrafast is being provided through OpenAI’s partnership with AI chip company Cerebras. Its hardware is designed around wafer-scale processors built to accelerate AI workloads, and the companies have previously worked together on high-speed inference products.

Cerebras said that GPT-5.6 Sol, running in Ultrafast mode, can reach up to 750 output tokens per second. The company is positioning faster inference as a way to make more advanced models practical for interactive applications where users cannot wait for lengthy responses.

OpenAI is initially limiting Ultrafast access to a select group of customers during the preview. The company said availability will expand as additional capacity becomes available.

The release adds another performance option for developers deploying GPT-5.6 Sol. Instead of reducing model capabilities to gain speed, OpenAI is testing whether specialised infrastructure can make its most capable models responsive enough for workloads that require near-real-time output.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Angry Angry 0
Sad Sad 0
Wow Wow 0
Shivangi Yadav Shivangi Yadav reports on startups, technology policy, and other significant technology-focused developments in India for TechAmerica.Ai. She previously worked as a research intern at ORF.