# Cerebras says GPT-5.6 Sol Ultrafast reaches 750 output tokens per second

> Source: <https://www.snipvote.com/story/cmssmhrnm000cye5ysiro3mlg>
> Published: 2026-08-14 07:35:21.272669+00:00

[Hacker News](https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai)

### Cerebras says GPT-5.6 Sol Ultrafast reaches 750 output tokens per second

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Cerebras and OpenAI launched GPT-5.6 Sol Ultrafast, achieving 750 output tokens per second without compromising quality, enabling significant speedups in mission-critical applications. This resolves the tradeoff between speed and intelligence, allowing for frontier AI models to be used in latency-sensitive tasks. The new service is initially available to a select group of customers.

Cerebras is now serving a frontier OpenAI model at up to 750 output tokens/sec with no reported accuracy loss, delivering roughly 5-7x end-to-end speedups on real reasoning workloads (HLE, GDP-Val) versus fast-mode competitors. This collapses the long-standing speed-vs-intelligence tradeoff for agentic loops, making it viable to put high-reasoning models directly on the critical path for latency-sensitive work like incident response, security triage, and interactive coding—but it's currently gated to a select customer set, so plan around limited availability and likely premium pricing.

### AI vs. AI Debate

“The summary could note that the speedup is achieved through Cerebras' specific hardware and partnership with OpenAI, which may impact the generalizability and cost structure of the Ultrafast service.”

“My summary explicitly flags "limited availability and likely premium pricing," directly addressing the cost and access implications that stem from the Cerebras-specific hardware partnership.”
