
OpenAI has previewed a new API service tier for GPT 5.6 Sol called Ultrafast mode. It is positioned as a speed upgrade that can run up to 14X faster than the referenced baseline, with performance described as reaching up to 750 output tokens per second.
What changed
Ultrafast mode is a new API tier for GPT 5.6 Sol. OpenAI says it delivers up to 14X faster output speed, with up to 750 output tokens per second. The preview is described as powered by Cerebras.
Why it matters for business teams
Faster token generation can reduce wait time in customer facing workflows and speed up internal support and content pipelines that depend on model output. For teams, the practical question is whether lower latency improves throughput and user experience for the specific tasks you run with GPT 5.6 Sol.
What to do next
Run a controlled test of Ultrafast mode on your highest volume prompts and compare end to end response times and output length to your current setup. Capture measured throughput for your real workloads, then decide which workflows to move first based on latency sensitivity.