cd /news/artificial-intelligence/openai-previews-ultrafast-api-tier-f… · home topics artificial-intelligence article
[ARTICLE · art-95952] src=testingcatalog.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

OpenAI previews Ultrafast API tier for GPT-5.6 Sol

OpenAI has opened a limited preview of Ultrafast, a new API service tier that runs GPT-5.6 Sol at up to 14 times the speed of Standard processing, generating up to 750 output tokens per second. Powered by Cerebras, the tier is initially restricted to select customers, with early testers including Jane Street, Podium, Basis, and Rogo. OpenAI says the tier is designed for low-latency applications such as incident response, finance, and customer support, and has not announced pricing or a general-availability timeline.

read2 min views1 publishedAug 13, 2026
OpenAI previews Ultrafast API tier for GPT-5.6 Sol
Image: Testingcatalog (auto-discovered)

OpenAI has opened a limited preview of Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol at up to 14 times the speed of Standard processing. Powered by Cerebras, the mode can generate up to 750 output tokens per second. Access is initially restricted to a select group of customers, with a wider rollout planned as capacity grows.

The service is designed for products and workflows where delays can determine whether an answer is still useful. OpenAI says achieving real-time speeds has often required teams to choose a smaller or specialized model. Ultrafast instead puts the company's most intelligent model into low-latency settings, seeking to deliver more useful work per second without making that tradeoff.

The immediate targets span incident response, finance, security, customer support, voice, commerce, and research. Teams could analyze logs, code changes, transactions, or market signals while events are unfolding. Voice and support systems could resolve multi-step requests without breaking a conversation, while commerce tools could check inventory, tailor recommendations, and address checkout problems before a shopper leaves. Researchers could also test and adjust work in shorter cycles.

OpenAI is already using the tier internally. During incidents, its developers have applied it to logs, traces, team conversations, follow-up checks, and preparation or validation of fixes, while engineers retain responsibility for judgment and deployment. Research teams are using it across connected tools to search knowledge sources, query data, and organize findings. OpenAI says some experiment loops that once ran overnight can now support several iterations within a workday.

Early testing includes Jane Street, Podium, Basis, and Rogo. Their feedback centers on more focused coding sessions, faster complex voice calls, low-latency applications built around a frontier model, and financial research that feels closer to a live exchange.

Ultrafast extends OpenAI's partnership with Cerebras, whose infrastructure supports GPT-5.6 Sol at the stated output rate. The preview is available through the API, and OpenAI is collecting input from initial customers to guide the service as capacity expands. A sign-up form is available for access updates, but the company has not announced pricing or a general-availability timeline.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-previews-ultr…] indexed:0 read:2min 2026-08-13 ·