cd /news/artificial-intelligence/openai-previews-cerebras-powered-gpt… · home topics artificial-intelligence article
[ARTICLE · art-96670] src=mlq.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

OpenAI previews Cerebras-powered GPT-5.6 Sol tier at up to 750 tokens per second

OpenAI previewed a Cerebras-powered service tier for its GPT-5.6 Sol model that generates up to 750 output tokens per second, claiming it is up to 14 times faster than its Standard processing tier. The limited API preview is available to select customers, with access expanding as Cerebras capacity grows, and OpenAI did not disclose pricing, regions, or an uptime SLA for the Ultrafast tier. The launch builds on a multiyear Cerebras agreement for up to 750 megawatts of inference capacity, reported by Reuters at more than $10 billion and later described by Cerebras as worth more than $20 billion.

read5 min views1 publishedAug 14, 2026
OpenAI previews Cerebras-powered GPT-5.6 Sol tier at up to 750 tokens per second
Image: Mlq (auto-discovered)
  • OpenAI says Ultrafast can generate up to 750 output tokens per second and is up to 14 times faster than Standard processing. [1] - The service is a limited API preview for selected customers, with access expanding as Cerebras capacity grows. The launch announcement does not state an Ultrafast price, supported regions or an uptime SLA. [1] - The launch builds on a multiyear Cerebras agreement for up to 750 megawatts of inference capacity. Reuters reported the January agreement at more than $10 billion; Cerebras later described it as worth more than $20 billion. [3][4] - OpenAI is testing interactive workloads including incident response, voice support, coding, financial research, commerce and live experimentation.

[1] OpenAI is previewing a Cerebras-powered service tier that runs its GPT-5.6 Sol model at up to 750 output tokens per second, the company said Thursday, August 13, 2026. OpenAI describes Ultrafast as up to 14 times faster than its Standard processing tier, positioning the service for applications where a frontier model must respond during a live interaction rather than after a longer batch job. [1]

The service is limited to a select group of API customers. OpenAI said it will expand access as capacity grows. The company’s announcement did not provide a separate Ultrafast price, regional availability list or uptime commitment. [1]

A speed claim tied to one model and one service tier #

OpenAI’s comparison is specific: Ultrafast runs GPT-5.6 Sol, the flagship model in the GPT-5.6 family, and measures performance against Standard processing. The company gives a peak generation rate of 750 output tokens per second and says the tier is up to 14 times faster than Standard. [1]

That is different from OpenAI’s existing Fast mode, which is available on a pay-as-you-go basis and promises up to 2.5 times faster speeds for GPT-5.6 Sol. OpenAI lists Fast mode at $10 per 1 million input tokens and $60 per 1 million output tokens, compared with standard Sol pricing of $5 and $30. Fast mode’s published table includes a 99.9% uptime SLA and a latency target for GPT-5.6 Sol, but those terms are identified as applying to enterprise customers. OpenAI has not said that the Fast mode pricing or guarantees apply to Ultrafast. [2][5]

The 14-times figure should therefore be read as an OpenAI estimate for its own service configurations, not as a direct independent comparison with another provider. OpenAI’s GPT-5.6 preview materials say its latency and cost estimates are based on simulated production behavior and can vary with tool calls, sampled tokens and input length. [6] Cerebras has separately said that Sol on its systems can run at 750 tokens per second and described the speed advantage over regular Sol as up to 10 times in a July guide, using a different comparison point from the new OpenAI announcement.

[7]## The commercial structure is still only partly public Ultrafast is the latest product layer in a broader infrastructure relationship. In January, OpenAI and Cerebras announced a multiyear agreement to deploy up to 750 megawatts of Cerebras inference systems in stages through 2028. Reuters reported that the deal was worth more than $10 billion, citing a source familiar with the matter. [3] Cerebras later disclosed in its first-quarter results announcement that it valued the agreement at more than $20 billion.

[4]The companies have not disclosed how much of that contracted capacity is supporting Ultrafast, how capacity is allocated among OpenAI products, or whether customers will access Cerebras systems directly. OpenAI’s announcement describes Ultrafast as an OpenAI API service tier, while Cerebras markets its own inference platform separately with pay-per-token developer access and enterprise offerings that include dedicated queue priority and uptime guarantees. [1][8]

The distinction matters for customers assessing cost and operational risk. Cerebras publishes a general claim that its inference can run in-region, but OpenAI’s Ultrafast announcement does not identify deployment regions, data-residency options, capacity quotas or reliability targets for this preview. [1][9]

OpenAI is testing real-time workflows first #

OpenAI said early Ultrafast customers include Jane Street, Podium, Basis and Rogo, with testing across coding, commerce, financial research, customer support and other interactive applications. The named companies described uses involving developer assistance, voice interactions and financial research. [1]

OpenAI also said its own engineers are using the service for incident response: reading logs and traces, synthesizing reports, checking hypotheses and helping prepare or validate fixes while an outage is in progress. A separate internal research workflow uses faster searches, data queries and experiment loops that can be repeated during the workday instead of being reviewed the next morning. [1]

Those examples describe intended operating conditions rather than evidence of broad production availability. OpenAI calls the offering an early look and says the preview is designed to identify where a large speed increase changes the value of a product. The company has not published aggregate customer latency, error-rate or availability data for Ultrafast. [1]

Companies mentioned #

Further sources #

[[1] OpenAI’s August 13, 2026 announcement describes Ultrafast, its 750-output-token… ↗](https://openai.com/index/previewing-ultrafast/)

[[2] OpenAI’s Fast mode documentation lists the existing service’s speed claim, pric… ↗](https://openai.com/api-fast-mode/)

[3] Reuters reported the January 2026 OpenAI-Cerebras agreement for up to 750 megaw… ↗

[4] Cerebras’ June 2026 results announcement described the OpenAI agreement as a mu… ↗

[[5] OpenAI’s Fast mode documentation states that its published uptime and latency S… ↗](https://openai.com/api-fast-mode/)

[[6] OpenAI’s GPT-5.6 preview materials explain that latency and cost estimates are … ↗](https://openai.com/index/previewing-gpt-5-6-sol/)+3 more

The stories that matter, in one email. Free — unsubscribe anytime.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-previews-cere…] indexed:0 read:5min 2026-08-14 ·