cd /news/artificial-intelligence/amd-s-mi355x-undercuts-nvidia-s-b300… · home topics artificial-intelligence article
[ARTICLE · art-83477] src=startupfortune.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

AMD's MI355X Undercuts Nvidia's B300 on Cost to Run China's Kimi K3

AMD's Instinct MI355X GPUs now run Moonshot AI's 2.8-trillion-parameter Kimi K3 model at 952 tokens per second per node, undercutting Nvidia's B300 on cost per token despite Nvidia's 1.65x raw throughput advantage, according to SemiAnalysis's InferenceX benchmarks. The MI355X rents at $2.50 per GPU-hour versus $6.00 for the B300, and AMD's chip delivers over 3.8 times the aggregate throughput of a B200 node on the same workload. AMD announced Day 0 inference support on July 27, 2026, with Moonshot AI, SGLang, and vLLM teams, positioning AMD as a cost-effective alternative for large-scale inference.

read4 min views2 publishedAug 2, 2026
AMD's MI355X Undercuts Nvidia's B300 on Cost to Run China's Kimi K3
Image: Startupfortune (auto-discovered)

AMD's Instinct MI355X now runs Moonshot AI's giant Kimi K3 model for a fraction of what Nvidia's B300 costs, and the benchmark numbers are public.

On July 27, 2026, AMD announced Day 0 inference support for Kimi K3 on its Instinct MI355X GPUs, working alongside Moonshot AI and the SGLang and vLLM teams to get the model running the moment its weights shipped. Kimi K3 is not a small model. It has 2.8 trillion parameters, making it roughly 75% larger than DeepSeek's V4 Pro, and Moonshot AI has called it the largest open-weight model released to date. Getting a model that size to run at all on day one is an engineering feat. Getting it to run cheaply is the story that actually matters to buyers.

That's where the MI355X pulled ahead. According to benchmarking data from SemiAnalysis's InferenceX platform, Nvidia's B300 still wins on raw aggregate throughput: it delivers about 1.65 times the output of a comparable MI355X deployment. Nvidia is still faster. But the B300 costs 2.4 times more to rent, at roughly $6.00 per GPU-hour against $2.50 for the MI355X. Do the math and the MI355X wins on the number that actually shows up on a cloud bill: cost per token generated. Cheaper wins. In real inference tests, AMD's chip hit 952 tokens per second per node. That's more than 3.8 times the aggregate throughput of a comparable Nvidia B200 node running the same workload.

The advantage traces to memory, not raw compute. Each chip carries 288 GiB of HBM3E and 8 TB/s of bandwidth, enough that Kimi K3 can run in an eight-GPU tensor-parallel configuration without the model spilling awkwardly across nodes. That matters, because Kimi K3 is enormous. Artificial Analysis noted that Kimi K3 requires roughly 1.56 TB of memory for its weights alone at 4-bit precision, a footprint too large for any single Hopper, B200, or even MI300X node. Only the newest chips clear that bar: Nvidia's B300 and AMD's MI350X and MI355X. Two products, then. And on this workload, AMD's is cheaper.

The market got there first #

The timing helps. AMD shares jumped 13% on July 30 after Microsoft reported that Azure crossed $100 billion in annual revenue for the first time, with growth accelerating to 43% as Microsoft's full-year capital spending hit nearly $116 billion, according to Microsoft's fiscal Q4 2026 earnings report covered by Fortune and 24/7 Wall St. That rally was built on a narrative: that hyperscalers are still spending aggressively on AI infrastructure, and that AMD stands to capture more of it. Kimi K3 gives that narrative a concrete data point instead of a promise. It's a real frontier model, independently benchmarked, running cheaper on AMD silicon than on Nvidia's newest chip.

Nvidia isn't losing outright here. The B300 remains the faster chip in absolute terms, and for workloads where raw latency matters more than cost per token, that 1.65x throughput edge still counts. Nvidia also retains the vastly larger installed base and software ecosystem that keeps most large training runs on its hardware. Training still belongs to Nvidia. But inference, not training, is where the AI industry's actual operating costs live once a model ships, and inference is exactly where this benchmark says AMD now competes on economics rather than just spec sheets.

More than a benchmark chart #

There's a geopolitical layer here too. Moonshot AI is a Chinese lab operating under U.S. export restrictions on advanced Nvidia chips sold into China, and its decision to prioritize AMD hardware for a flagship release is worth watching, though nothing in AMD's or Moonshot's public materials frames it as a deliberate snub of Nvidia. What's clear is who did the work: the deployment was a joint effort involving Moonshot AI, AMD, Nvidia, the RadixArk-based SGLang and Miles team, Baseten, and Modal, with vLLM and SGLang both shipping day-0 support across Nvidia's Hopper and Blackwell chips and AMD's MI355X simultaneously. Frankly, that's the more interesting detail than any single benchmark chart: the entire inference stack, from Chinese model to American chipmaker to open-source serving frameworks, moved in lockstep on day one, regardless of which country's export rules apply to which GPU.

Kimi K3 also isn't a niche release padding out a benchmark. VentureBeat reported that the model beat Claude Opus 4.8 and GPT-5.5 on Moonshot's own coding and agentic evaluation suite. A model that competitive, running on hardware that's demonstrably cheaper per token, is exactly the kind of proof point AMD needed after years of being told its inference software stack couldn't keep pace with Nvidia's CUDA ecosystem. It hasn't closed that gap everywhere. On this one, with real weights and a real published benchmark, it has.

Also read: Anthropic Admits Its Own Bugs Broke Claude Code After Weeks of DenialAmazon Shuts Its AGI Lab and Cuts Jobs to Chase Enterprise AI InsteadNXP is chasing Ambarella for $3.3 billion because it needs eyes, not just brains

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @amd 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/amd-s-mi355x-undercu…] indexed:0 read:4min 2026-08-02 ·