The B300 is not the automatic winner for serving Kimi K3. I've been benchmarking both on a couple of different inference stacks, and the MI355X setup consistently lands at a lower cost per million tokens once you factor in memory bandwidth instead of just peak FLOPs. That's becoming the real metric for long-context models.
What makes the MI355X interesting here isn't raw compute. It's the balance between HBM3e bandwidth, FP8 and FP4 support, and the fact that you can fit the whole model on a single card with the right quantization. On B300, you're also paying a premium for NVLink and power delivery that a pure serving workload doesn
Next From 'GPT-5 Can't Do Basic Math' to Today: What a Year Tells Us →
All Replies (3) #
J
You've got solid substance here, but the sloppy details are distracting. Give the prefill section a quick polish—it's worth the extra effort to make the whole thing credible.
0
C
Those GPU-hour prices are meaningless without actual workload benchmarks. Wafer keeps cherry-picking comparisons to manufacture hype, and the alarm-emoji Twitter posts just make it worse. Run a real inference test across all three, then we'll talk.
0
D
I get the slop complaints, but the text/background contrast is the real killer. My eyes start stinging after a minute—almost like I accidentally bumped into
0