Reddit
OpenAI Previews 'Ultrafast' Tier: GPT-5.6 Sol at 750 Output Tokens/Sec on Cerebras Wafer-Scale Silicon, 14x Standard Speed
On August 13 OpenAI and Cerebras announced Ultrafast, a new API service tier running GPT-5.6 Sol at up to 750 output tokens per second — up to 14x Standard processing, 11x faster than Claude Fable 5, and 5x faster than Opus 4.8 on Fast mode. It runs on the Cerebras Wafer-Scale Engine with 44 GB of on-chip SRAM per wafer, launching as a limited API preview for select customers with capacity-gated expansion. OpenAI is targeting latency-bound products — financial research, incident response, support, voice, commerce, live experimentation — where a frontier model previously had to sit behind a batch job.
Source
↳ Follow the thread