The programming language speed balls, but for inference (tokens per second) Cerebras' gpt-oss-120B (low) leads a new inference speed benchmark at 1,538.0 tokens per second, outpacing Groq's gpt-oss-20B (high) at 949.5 t/s and Cerebras' GPT-5.6 Sol (ultrafast preview) at up to 750 t/s, according to a live leaderboard tracking tokens per second for major AI models. The ranking includes 18 models from providers such as Google AI Studio, OpenAI, Anthropic, and Alibaba Cloud, with speeds ranging from 1,538.0 t/s down to 47.0 t/s for Qwen3.8 Max. See the balls See the actual tokens Ⅱ Pause ↻ Reset 1,000 output tokens One trip to the other side = 1,000 output tokens · A round trip = 2,000 gpt-oss-120B low Cerebras · 1,538.0 t/s 120B params ↗ gpt-oss-20B high Groq · 949.5 t/s 20B params ↗ GPT-5.6 Sol ultrafast preview Cerebras · up to 750 t/s Undisclosed params ↗ Nemotron 3.5 Lightning DeepInfra NVFP4 · 516.2 t/s 31.6B / 3.6B params ↗ Gemini 3.7 Flash high Google AI Studio · 340.1 t/s Undisclosed params ↗ DeepSeek V4 Flash Non-reasoning Makora · 274.1 t/s 284B / 13B params ↗ MiniMax-M3 Modular · 231.5 t/s 428B / 23B params ↗ Muse Spark 1.2 xhigh Meta · ~217.2 t/s Undisclosed params ↗ GLM-5.2 max Baseten FAST · 195.5 t/s 744B params ↗ Gemini 3.5 Flash high Google AI Studio · 172.9 t/s Undisclosed params ↗ GPT-5.6 Luna max OpenAI · 150.2 t/s Undisclosed params ↗ Claude Sonnet 5 Adaptive Reasoning, Max Effort Google · 81.8 t/s Undisclosed params ↗ Grok 4.6 high SpaceXAI · 65.5 t/s Undisclosed params ↗ Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback Anthropic · 63.2 t/s Undisclosed params ↗ GPT-5.6 Sol max OpenAI · 61.7 t/s Undisclosed params ↗ Claude Opus 5 Adaptive Reasoning, Max Effort Anthropic · 54.1 t/s Undisclosed params ↗ Kimi K3 max Fireworks · 48.9 t/s 2.8T params ↗ Qwen3.8 Max Alibaba Cloud · 47.0 t/s Undisclosed params ↗