Skip to content
NewEraAI

AI news

Ultrafast mode for GPT 5.6 Sol promises up to 14X faster API responses

A new OpenAI API service tier called Ultrafast mode is previewed for GPT 5.6 Sol, claiming up to 14X faster generation speed. The tier is described as producing up to 750 output tokens per second.

13 August 2026

Close-up of a computer screen displaying ChatGPT interface in a dark setting.
Photograph by Matheus Bertelli · Pexels

OpenAI has previewed a new API service tier for GPT 5.6 Sol called Ultrafast mode. It is positioned as a speed upgrade that can run up to 14X faster than the referenced baseline, with performance described as reaching up to 750 output tokens per second.

What changed

Ultrafast mode is a new API tier for GPT 5.6 Sol. OpenAI says it delivers up to 14X faster output speed, with up to 750 output tokens per second. The preview is described as powered by Cerebras.

Why it matters for business teams

Faster token generation can reduce wait time in customer facing workflows and speed up internal support and content pipelines that depend on model output. For teams, the practical question is whether lower latency improves throughput and user experience for the specific tasks you run with GPT 5.6 Sol.

What to do next

Run a controlled test of Ultrafast mode on your highest volume prompts and compare end to end response times and output length to your current setup. Capture measured throughput for your real workloads, then decide which workflows to move first based on latency sensitivity.

Next step

Start with the free AI Opportunity Assessment.

A short, no-obligation conversation about where enquiries, hours and revenue leak today. You do not have to pick a tier to have it, and what comes out of it feeds Discover, so the first paid day starts from evidence rather than a blank sheet.