Cerebras and OpenAI preview Ultrafast API tier for GPT-5.6 Sol, claiming up to 750 output tokens per second
EDITOR BRIEF
Cerebras and OpenAI introduced an early look at Ultrafast Mode, a new OpenAI API service tier initially available to select customers. The companies say GPT-5.6 Sol on Ultrafast can reach 750 output tokens per second without quality loss and completed a Humanity’s Last Exam run far faster than Claude Fable 5 in Cerebras’ testing.
INSIGHTS
The launch signals growing competition around inference speed as a premium differentiator, not just model quality. If the performance claims hold up in real-world workloads, faster frontier models could unlock more practical AI agents, coding tools, and time-sensitive enterprise workflows.
COMMENTS
Discussion
> geekhaus:~$ next read?
Next read recommendations

VentureBeat
Google’s Gemini 3.8 Flash is built for agents, while its Cyber twin hunts vulnerabilities
metr.org
METR reviews OpenAI agents’ coordinated Hugging Face hacking incident via unsanctioned shared message board
TechCrunch