{"slug": "googles-tpu-v7-now-beats-nvidias-gb200-by-56-on-kimi-k3-and-is-closing-in-on-on", "title": "Google’s TPU v7 Now Beats Nvidia’s GB200 By 56% On Kimi K3, And Is Closing In On Cerebras On Qwen 3.8, Says Semi Analysis", "summary": "Inferact, a startup whose team includes vLLM maintainers, released open-source \"megakernels\" for Google's TPU v7 (Ironwood) that delivered 709 tokens per second for a single user on Moonshot AI's Kimi K3 across 16 chips, versus 452 tokens per second for 16 Nvidia GB200 GPUs running vLLM's published recipe — roughly 1.6x, or 56% better, according to SemiAnalysis. On Alibaba's Qwen 3.8 27B, four TPU v7 chips reached 1,515 tokens per second per user against 695 on four GB200s, coming \"close to\" Cerebras's approximately 1,850 tokens per second per user on the same model. The results rely on speculative decoding with the DSpark method at an acceptance length of six, and Inferact reported 0.944 on GPQA-Diamond and 0.972 on GSM8K to show accuracy was not degraded.", "body_md": "Google’s TPU v7 has just posted some of the most striking inference numbers yet outside Nvidia’s ecosystem, thanks to some heavily optimised software from the team behind vLLM’s commercial platform.\n\nInferact, a startup whose team includes vLLM maintainers, has released a set of open-source “megakernels” for TPU v7, also known as Ironwood. On Moonshot AI’s Kimi K3, 16 TPU v7 chips delivered 709 tokens per second for a single user, versus 452 tokens per second for 16 Nvidia GB200 GPUs running vLLM’s published recipe. That works out to roughly 1.6x the throughput, or 56% better, as SemiAnalysis put it.\n\nThe firm also flagged a second result. On Alibaba’s Qwen 3.8 27B, just four TPU v7 chips reached 1,515 tokens per second per user, compared with 695 on four GB200s, according to Inferact’s data. SemiAnalysis’s chart put Cerebras at roughly 1,850 tokens per second per user on the same model, prompting the firm to say four TPUs had come “close to” a single giant Cerebras wafer on interactivity.\n\n## What’s a megakernel?\n\nDecoding, the phase where a model generates tokens one after another, is limited mostly by how quickly weights can be moved from memory to the chip. In a conventional setup, a single decode step launches hundreds of separate kernels, each of which loads data, computes, and writes results back. Every handoff leaves gaps where memory bandwidth sits idle, and those gaps add up.\n\nA megakernel folds the whole decoder step into one program. That lets the software start fetching the next layer’s weights while the current layer is still computing, keeping the memory pipeline busy throughout. Inferact’s Kimi K3 kernel packs all 92 of the model’s mixture-of-experts layers into a single call, and is spread across 16 chips with 32 TensorCores.\n\n## Why TPUs suit the approach\n\nOn paper, TPU v7 and the GB200 are close. Inferact lists 2.31 petaflops of BF16 compute and 7.38 TB/s of memory bandwidth for the TPU, against 2.5 petaflops and 8 TB/s for the Nvidia part. The difference lies in on-chip memory. Each TPU TensorCore has 64 MiB of software-managed on-chip memory called VMEM, large enough to stage most of a layer’s weights ahead of time. An entire GB200 GPU, by comparison, has around 38 MiB of comparable tensor memory, divided among 152 streaming multiprocessors.\n\nThe TPU’s sequential programming model also helped. Inferact wrote the kernel in Pallas, Google’s kernel language, and says it’s the first open-source inference megakernel built with it. As a side benefit, the full kernel compiles from scratch in under 90 seconds, whereas large TPU models built from hundreds or thousands of XLA operations can take over 30 minutes.\n\n## The numbers, with caveats\n\nThe headline 709 tokens per second figure relies on speculative decoding, where a small draft model proposes several tokens that the main model verifies in a single step. Inferact’s setup used the DSpark method with an acceptance length of six, meaning six tokens were accepted per step on average. At an acceptance length of three, the gap was smaller, at 350 versus 229 tokens per second. Each decode step took about 8.5 milliseconds.\n\nWithout speculative decoding, Inferact says its kernels for Kimi K3 and Qwen 3.8 27B deliver roughly 1.4x to 2x the decode throughput of the GB200 baseline at batch sizes from one to eight, nearly double at batch size one. It reported scores of 0.944 on GPQA-Diamond and 0.972 on GSM8K to show the optimisations hadn’t hurt accuracy.\n\nA few things are worth keeping in mind. The GB200 baseline is vLLM’s stock recipe on 16 GPUs, not a bespoke hand-tuned kernel, and not a full NVL72 rack. The results are for low-batch, single-user speed, which is where interactivity matters most for agents and coding tools, not for maximum throughput across many users. And the Cerebras figure in SemiAnalysis’s chart is marked approximate, with no methodology attached. Cerebras, which has become one of the more prominent AI hardware names, thanks in part to its [multi-billion-dollar deal with OpenAI](https://officechai.com/ai/sam-altman-greg-brockman-didnt-inform-musk-of-their-personal-investments-in-cerebras-while-openai-was-looking-to-acquire-it-in-2017/), hasn’t yet responded to the comparison in the material we’ve seen.\n\n## The bigger picture: TPUs going external\n\nSemiAnalysis framed the results as part of a broader push by Google to make its TPUs usable beyond its own walls. The firm said Google’s [“externalization of software”](https://officechai.com/ai/googles-tpuv7-ironwood-achieves-50-better-performance-per-dollar-than-nvidias-blackwell-ultra-says-semi-analysis/) around TPUs is “full steam ahead”, and that it’s important to follow. SemiAnalysis had earlier found Ironwood delivering up to 50% better performance per dollar than Nvidia’s Blackwell Ultra, and TPU adoption has grown among labs looking for alternatives to Nvidia’s chips.\n\nThe workloads chosen matter too. [Kimi K3](https://officechai.com/ai/kimi-k3-2-8-trillion-parameters-pricing-context-window/) is a 2.8-trillion-parameter open-weight model with a one-million-token context window, and Moonshot has said it ranks behind only Claude Fable 5 and GPT-5.6 Sol on overall intelligence. Its [weights are now public](https://officechai.com/ai/moonshot-ai-releases-kimi-k3s-weights-sees-fastest-release-growth-ever-on-hugging-face/), which means anyone can serve it and compete on speed. The 27-billion-parameter [Qwen 3.8](https://officechai.com/miscellaneous/alibaba-releases-qwen-3-8-27b-beats-muse-glimmer-30b-on-many-benchmarks/) model, meanwhile, is pitched at builders who want strong coding and agentic performance in a compact package. Fast serving of both matters for the agentic workloads where a model makes many sequential calls and latency compounds.\n\nInferact says it’s just getting started. Its roadmap includes higher-concurrency workloads, agentic use cases where KV cache movement becomes the bottleneck, and topologies beyond the 16-chip 2x2x4 layout the current kernel is tailored to. The code is available on GitHub as `inferact/tpu-megakernels`.", "url": "https://wpnews.pro/news/googles-tpu-v7-now-beats-nvidias-gb200-by-56-on-kimi-k3-and-is-closing-in-on-on", "canonical_source": "https://officechai.com/ai/googles-tpu-v7-now-beats-nvidias-gb200-by-56-on-kimi-k3-and-is-closing-in-on-cerebras-on-qwen-3-8-says-semi-analysis/", "published_at": "2026-09-28 11:25:38+00:00", "updated_at": "2026-09-28 11:48:20.587618+00:00", "lang": "en", "topics": ["ai-chips", "ai-infrastructure", "large-language-models", "ai-research"], "entities": ["Google", "TPU v7", "Nvidia", "GB200", "Inferact", "vLLM", "SemiAnalysis", "Cerebras"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/googles-tpu-v7-now-beats-nvidias-gb200-by-56-on-kimi-k3-and-is-closing-in-on-on", "markdown": "https://wpnews.pro/news/googles-tpu-v7-now-beats-nvidias-gb200-by-56-on-kimi-k3-and-is-closing-in-on-on.md", "text": "https://wpnews.pro/news/googles-tpu-v7-now-beats-nvidias-gb200-by-56-on-kimi-k3-and-is-closing-in-on-on.txt", "jsonld": "https://wpnews.pro/news/googles-tpu-v7-now-beats-nvidias-gb200-by-56-on-kimi-k3-and-is-closing-in-on-on.jsonld"}}