cd /news/artificial-intelligence/the-programming-language-speed-balls… · home topics artificial-intelligence article
[ARTICLE · art-96139] src=kylejeong.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

The programming language speed balls, but for inference (tokens per second)

Cerebras' gpt-oss-120B (low) leads a new inference speed benchmark at 1,538.0 tokens per second, outpacing Groq's gpt-oss-20B (high) at 949.5 t/s and Cerebras' GPT-5.6 Sol (ultrafast preview) at up to 750 t/s, according to a live leaderboard tracking tokens per second for major AI models. The ranking includes 18 models from providers such as Google AI Studio, OpenAI, Anthropic, and Alibaba Cloud, with speeds ranging from 1,538.0 t/s down to 47.0 t/s for Qwen3.8 Max.

read1 min views1 publishedAug 14, 2026

See the balls See the actual tokens Ⅱ

↻ Reset 1,000 output tokens One trip to the other side = 1,000 output tokens · A round trip = 2,000 gpt-oss-120B (low) Cerebras · 1,538.0 t/s 120B params ↗ gpt-oss-20B (high) Groq · 949.5 t/s 20B params ↗ GPT-5.6 Sol (ultrafast preview) Cerebras · up to 750 t/s Undisclosed params ↗ Nemotron 3.5 Lightning DeepInfra (NVFP4) · 516.2 t/s 31.6B / 3.6B params ↗ Gemini 3.7 Flash (high) Google AI Studio · 340.1 t/s Undisclosed params ↗ DeepSeek V4 Flash (Non-reasoning) Makora · 274.1 t/s 284B / 13B params ↗ MiniMax-M3 Modular · 231.5 t/s 428B / 23B params ↗ Muse Spark 1.2 (xhigh) Meta · ~217.2 t/s Undisclosed params ↗

GLM-5.2 (max)
Baseten (FAST) · 195.5 t/s

744B params ↗ Gemini 3.5 Flash (high) Google AI Studio · 172.9 t/s Undisclosed params ↗ GPT-5.6 Luna (max) OpenAI · 150.2 t/s Undisclosed params ↗ Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Google · 81.8 t/s Undisclosed params ↗ Grok 4.6 (high) SpaceXAI · 65.5 t/s Undisclosed params ↗ Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Anthropic · 63.2 t/s Undisclosed params ↗ GPT-5.6 Sol (max) OpenAI · 61.7 t/s Undisclosed params ↗ Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic · 54.1 t/s Undisclosed params ↗ Kimi K3 (max) Fireworks · 48.9 t/s 2.8T params ↗ Qwen3.8 Max Alibaba Cloud · 47.0 t/s Undisclosed params ↗

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @cerebras 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-programming-lang…] indexed:0 read:1min 2026-08-14 ·