AMD's Instinct MI355X now runs Moonshot AI's giant Kimi K3 model for a fraction of what Nvidia's B300 costs, and the benchmark numbers are public.
On July 27, 2026, AMD announced Day 0 inference support for Kimi K3 on its Instinct MI355X GPUs, working alongside Moonshot AI and the SGLang and vLLM teams to get the model running the moment its weights shipped. Kimi K3 is not a small model. It has 2.8 trillion parameters, making it roughly 75% larger than DeepSeek's V4 Pro, and Moonshot AI has called it the largest open-weight model released to date. Getting a model that size to run at all on day one is an engineering feat. Getting it to run cheaply is the story that actually matters to buyers.
That's where the MI355X pulled ahead. According to benchmarking data from SemiAnalysis's InferenceX platform, Nvidia's B300 still wins on raw aggregate throughput: it delivers about 1.65 times the output of a comparable MI355X deployment. Nvidia is still faster. But the B300 costs 2.4 times more to rent, at roughly $6.00 per GPU-hour against $2.50 for the MI355X. Do the math and the MI355X wins on the number that actually shows up on a cloud bill: cost per token generated. Cheaper wins. In real inference tests, AMD's chip hit 952 tokens per second per node. That's more than 3.8 times the aggregate throughput of a comparable Nvidia B200 node running the same workload.
The advantage traces to memory, not raw compute. Each chip carries 288 GiB of HBM3E and 8 TB/s of bandwidth, enough that Kimi K3 can run in an eight-GPU tensor-parallel configuration without the model spilling awkwardly across nodes. That matters, because Kimi K3 is enormous. Artificial Analysis noted that Kimi K3 requires roughly 1.56 TB of memory for its weights alone at 4-bit precision, a footprint too large for any single Hopper, B200, or even MI300X node. Only the newest chips clear that bar: Nvidia's B300 and AMD's MI350X and MI355X. Two products, then. And on this workload, AMD's is cheaper.
The market got there first #
The timing helps. AMD shares jumped 13% on July 30 after Microsoft reported that Azure crossed $100 billion in annual revenue for the first time, with growth accelerating to 43% as Microsoft's full-year capital spending hit nearly $116 billion, according to Microsoft's fiscal Q4 2026 earnings report covered by Fortune and 24/7 Wall St. That rally was built on a narrative: that hyperscalers are still spending aggressively on AI infrastructure, and that AMD stands to capture more of it. Kimi K3 gives that narrative a concrete data point instead of a promise. It's a real frontier model, independently benchmarked, running cheaper on AMD silicon than on Nvidia's newest chip.
Nvidia isn't losing outright here. The B300 remains the faster chip in absolute terms, and for workloads where raw latency matters more than cost per token, that 1.65x throughput edge still counts. Nvidia also retains the vastly larger installed base and software ecosystem that keeps most large training runs on its hardware. Training still belongs to Nvidia. But inference, not training, is where the AI industry's actual operating costs live once a model ships, and inference is exactly where this benchmark says AMD now competes on economics rather than just spec sheets.
More than a benchmark chart #
There's a geopolitical layer here too. Moonshot AI is a Chinese lab operating under U.S. export restrictions on advanced Nvidia chips sold into China, and its decision to prioritize AMD hardware for a flagship release is worth watching, though nothing in AMD's or Moonshot's public materials frames it as a deliberate snub of Nvidia. What's clear is who did the work: the deployment was a joint effort involving Moonshot AI, AMD, Nvidia, the RadixArk-based SGLang and Miles team, Baseten, and Modal, with vLLM and SGLang both shipping day-0 support across Nvidia's Hopper and Blackwell chips and AMD's MI355X simultaneously. Frankly, that's the more interesting detail than any single benchmark chart: the entire inference stack, from Chinese model to American chipmaker to open-source serving frameworks, moved in lockstep on day one, regardless of which country's export rules apply to which GPU.
Kimi K3 also isn't a niche release padding out a benchmark. VentureBeat reported that the model beat Claude Opus 4.8 and GPT-5.5 on Moonshot's own coding and agentic evaluation suite. A model that competitive, running on hardware that's demonstrably cheaper per token, is exactly the kind of proof point AMD needed after years of being told its inference software stack couldn't keep pace with Nvidia's CUDA ecosystem. It hasn't closed that gap everywhere. On this one, with real weights and a real published benchmark, it has.
Also read: Anthropic Admits Its Own Bugs Broke Claude Code After Weeks of Denial • Amazon Shuts Its AGI Lab and Cuts Jobs to Chase Enterprise AI Instead • NXP is chasing Ambarella for $3.3 billion because it needs eyes, not just brains