Twelve accelerators on memory per dollar, bandwidth per compute, MLPerf tokens per dollar, domain size, and software status.
Buying or renting AI compute in 2026 is a memory decision, not a FLOPS decision. This post puts twelve accelerators on the same axes using only vendor datasheets, public price lists, MLCommons result files, and SEC filings, checked on 9 September 2026. The derived metrics are ours. The inputs are not.
The chips #
| Chip | Vendor | Memory | Bandwidth | Dense FP8 | Scale-up domain | Status |
|---|---|---|---|---|---|---|
| H200 | NVIDIA | 141 GB HBM3e | 4.8 TB/s | 2.0 PFLOPS | 8 | Shipping, sold out in cloud per Q4 FY26 call |
| B200 (HGX) | NVIDIA | 180 GB HBM3e | 8.0 TB/s | 4.5 PFLOPS | 8 | Shipping |
| B300 (HGX) | NVIDIA | 270 GB HBM3e | 8.0 TB/s | 5.0 PFLOPS | 8 | Shipping |
| GB300 NVL72 | NVIDIA | 288 GB HBM3e | 8.0 TB/s | 5.0 PFLOPS | 72 | Shipping, most of NVIDIA volume |
| Vera Rubin NVL72 | NVIDIA | 288 GB HBM4 | 22 TB/s | Not disclosed | 72 | Production shipments began August 2026 |
| MI355X | AMD | 288 GB HBM3e | 8.0 TB/s | 5.03 PFLOPS | 8 | Shipping |
| MI455X (Helios) | AMD | 432 GB HBM4 | 23.3 TB/s | 20.1 PFLOPS | 72 | Ramp 2H 2026, no MLPerf, not in ROCm notes |
| Ironwood TPU7x | 192 GiB HBM | 7.38 TB/s | 4.61 PFLOPS | 9,216 pod | GA 31 March 2026, two zones | |
| Trainium3 | AWS | 144 GB HBM3e | 4.9 TB/s | 2.52 PFLOPS | 64 to 144 | GA December 2025, no public price |
| Blackhole p150 | Tenstorrent | 32 GB GDDR6 + 180 MB SRAM | 512 GB/s | 0.66 PFLOPS (block FP8) | Ethernet, 3.2 Tb/s per card | In stock, $1,399 |
| WSE-3 (CS-3) | Cerebras | 44 GB SRAM | 21 PB/s | 125 PFLOPS (FP16, vendor) | Wafer | Shipping, CS-4 shipments from Q3 2026 |
| Gaudi 3 | Intel | 128 GB HBM2e | 3.7 TB/s | 1.68 PFLOPS | 8 | Listed, written off in the FY2025 10-K |
Table 1. Vendor-stated per-chip figures. NVIDIA and AMD dense figures are from spec tables. Rubin FP8 is not published. Cerebras does not state precision for its 125 PFLOPS. Tenstorrent publishes block FP8 only.
Groq is absent from the chip table on purpose. Its next chip ships as the NVIDIA Groq 3 LPX under a licence, and Groq’s own newsroom describes GroqCloud as running NVIDIA hardware alongside LPUs. It is a cloud provider now, not an accelerator option you can buy.
Bandwidth per unit of compute #
Decode is bound by how fast weights and KV cache move, not by how fast the tensor cores multiply. The ratio below is HBM bandwidth in TB/s divided by dense FP8 PFLOPS. Lower means the chip is more compute-heavy relative to its memory, and spends more of its life waiting on HBM at inference batch sizes.
| Chip | TB/s per dense FP8 PFLOPS |
|---|---|
| H200 | 2.40 |
| Gaudi 3 | 2.20 |
| Trainium3 | 1.94 |
| B200 | 1.78 |
| Ironwood | 1.60 |
| B300 | 1.60 |
| MI355X | 1.59 |
| MI455X | 1.16 |
Table 2. Derived from Table 1. Every new generation lowers this ratio. FP8 and FP4 throughput grew faster than HBM bandwidth on every vendor’s roadmap.
The practical reading: a B300 has 2.5 times the FP8 of an H200 but only 1.67 times the bandwidth. For a 70B model at moderate batch, the bandwidth is what you are paying for. That makes memory per dollar the first sort key.
Memory per rented dollar #
Public on-demand prices are rare. These are the ones that exist, from the provider’s own price page or price API on the day.
| Chip | Provider | $/GPU-hour | GB per $/hr | TB/s per $/hr |
|---|---|---|---|---|
| MI355X | Vultr (bare metal, 8 GPU) | 2.59 | 111 | 3.09 |
| MI325X | Vultr (bare metal, 8 GPU) | 4.62 | 55 | 1.30 |
| Trainium2 | AWS Capacity Blocks | 2.24 | 43 | 1.30 |
| Ironwood | Google, 3-year commit | 5.40 | 36 | 1.37 |
| B300 | Nebius | 7.85 | 34 | 1.02 |
| MI355X | Oracle OCI | 8.60 | 33 | 0.93 |
| H200 | Nebius | 4.50 | 31 | 1.07 |
| B200 | Lambda | 6.69 | 27 | 1.20 |
| GB200 NVL72 | CoreWeave | 10.50 | 18 | 0.76 |
| Ironwood | Google, on demand | 12.00 | 16 | 0.62 |
Table 3. Prices as listed on 9 September 2026. Vultr’s MI355X price sits below its own MI325X price and may be promotional. GB300 NVL72, Trainium3, and Rubin have no public hourly price anywhere.
Two things stand out. AMD’s MI355X on Vultr delivers two and a half times the memory bandwidth per dollar of the best NVIDIA part with a public price. And Google’s on-demand Ironwood is the most expensive memory on the list, while the same chip on a three-year commitment is competitive. The commitment, not the silicon, is the product.
MLPerf tokens per dollar #
MLCommons publishes system totals. We divided each 8-GPU server result for Llama 2 70B by eight, then by the hourly price above.
Figure 1. Derived from MLPerf Inference v6.0 (B200, B300, MI355X) and v5.1 (H200, MI325X) server results divided by GPU count and price. Exact values in Table 4.
| Chip | MLPerf round | Tokens/s per GPU | $/hr | Tokens/s per $ |
|---|---|---|---|---|
| B300 | v6.0, NVIDIA DGX B300 | 13,414 | 7.85 | 1,709 |
| B200 | v6.0, HPE 8x B200 | 12,954 | 6.69 | 1,936 |
| MI355X | v6.0, AMD 8x MI355X | 12,535 | 8.60 (OCI) or 2.59 (Vultr) | 1,458 or 4,840 |
| GB300 NVL72 | v6.0, NVIDIA, per GPU of 72 | 12,059 | Not public | Not computable |
| H200 | v5.1, ASUS 8x H200 | 4,274 | 4.50 | 950 |
| MI325X | v5.1, AMD 8x MI325X | 4,003 | 4.62 | 867 |
Table 4. Per-GPU figures are our division of the published system total. Result IDs are in the MLCommons repositories for v5.1 and v6.0.
On this benchmark the MI355X is within 7 percent of the B300 per GPU. Price decides the rest. Ironwood, Trainium3, Gaudi 3, and Helios have no MLPerf inference submission in any round through v6.0, so they cannot be placed on this chart. Cerebras and Tenstorrent do not submit either.
Scale-up domain #
The number of accelerators that share memory at full interconnect speed sets the largest model you can serve without crossing a network hop.
| Platform | Domain | Per-accelerator link | Memory in domain |
|---|---|---|---|
| HGX B300, MI355X UBB, Gaudi 3 | 8 | 1.8 TB/s, 1.08 TB/s, 24x200 GbE | 2.1 TB, 2.3 TB, 1 TB |
| GB300 NVL72 | 72 | 1.8 TB/s NVLink 5 | 20 TB |
| Vera Rubin NVL72 | 72 | 3.6 TB/s NVLink 6 | 20.7 TB |
| Helios (MI455X) | 72 | 3.6 TB/s UALink over Ethernet | 31 TB |
| Trainium3 UltraServer | 64 or 144 | 2 TB/s NeuronLink-v4 | 9.2 TB or 20.7 TB |
| Ironwood pod | 9,216 | 1.2 TB/s ICI, 3D torus | 1.77 PB |
| Galaxy Blackhole | 32 chips, Ethernet | 800 GbE ports | 1 TB GDDR6 + 6.2 GB SRAM |
| CS-3 cluster | Up to 2,048 wafers | Fabric | MemoryX up to 1.2 PB |
Table 5. Vendor-stated. Helios matches Rubin on link bandwidth and exceeds it on memory, on paper. No customer shipment has been announced as of this writing.
Software status #
| Stack | State as of 9 September 2026 | Documented limitation |
|---|---|---|
| CUDA 13.3, TensorRT-LLM, Dynamo | Mature. Every MLPerf NVIDIA entry uses it | Rubin toolchain not yet in release notes |
| ROCm 10.0.0 (26 August 2026) | vLLM 0.27, SGLang 0.5.15, PyTorch 2.13 on MI355X | 9 to 25 percent training slowdown on gfx950 from an AOTriton kernel choice, workaround in notes. MI455X not in supported list |
| Neuron SDK 2.32 | Trainium3 supported since December 2025 | vLLM Neuron plugin is Beta. DeepSeek V3 and R1 are not production models |
| JAX and PyTorch on TPU7x | Supported | vLLM TPU plugin has no GA label. TensorFlow not supported |
| tt-metal, tt-forge (Apache 2.0) | Llama 3.1 8B “Complete” on Blackhole | Llama 3.3 70B is “Functional”, not “Complete”, on Blackhole. 70B numbers published for Wormhole Galaxy only: 72.5 tokens/s per user |
| Cerebras | Inference as API. gpt-oss-120b at $0.35 in, $0.75 out per million | Llama 3.3 70B removed from the public price list. No 2026 independent speed measurement found |
| Intel Gaudi 1.24 | vLLM plugin v0.26 upstream | Lazy mode deprecated. Intel’s 10-K calls the Gaudi effort “unsuccessful” |
Table 6. From release notes, model support matrices, and filings. “Complete” and “Functional” are Tenstorrent’s own support tiers.
What we would pick #
- Renting for 70B-class inference today. MI355X on Oracle or Vultr, or B200 on Lambda. Run the per-dollar table with your provider’s quote. The silicon gap is 7 percent. The price gap is up to 3x.
- Renting for models that need more than 2 TB in one domain. GB300 NVL72 is the only 72-GPU domain shipping in volume with a benchmark record. Rubin started shipping in August 2026 and has no MLPerf entry. Helios has neither shipments nor ROCm notes yet.
- Committed capacity for a year or more. Ironwood at the three-year rate is competitive on memory per dollar and comes with the largest domain in the industry. On demand it is the worst value here.
- Owning hardware under $10,000. Tenstorrent is the only option with an open stack and a public price. Treat 70B as experimental on Blackhole until the matrix says “Complete”.
- Lowest latency per user. Cerebras and Groq as APIs, not as hardware purchases. The Groq silicon roadmap now belongs to NVIDIA’s catalogue.
- Gaudi 3. No. Intel wrote off $1.3 billion of inventory across 2024 and 2025 and its 2026 earnings calls do not mention it.
Scope #
We measured nothing ourselves. Every performance number is a vendor datasheet or an MLCommons result file, and every price is a public list price on one day. We did not test training, and the bandwidth ratio in Table 2 is a heuristic for decode, not a benchmark. Cloud prices for GB300, Rubin, Trainium3, MI455X, and Groq 3 LPX do not exist publicly, so those chips cannot be ranked on cost. Tenstorrent’s Galaxy Blackhole list price moved from $110,000 in April 2026 to $160,000 in September 2026 on the vendor’s own pages. We will rerun the tables when MLPerf v6.1 lands.
Choosing accelerators for a build? Book a free 30 minute consultation.