Via intuitionlabs.ai
On-demand capacity is effectively sold out across providers, with even small clusters becoming difficult to secure
Renting Nvidia’s workhorse AI chip just got a lot more expensive. The cost of H100 GPU rentals climbed roughly 40% between October 2025 and March 2026, according to data from SemiAnalysis, with average rates jumping from $1.70 per GPU-hour to $2.35 per GPU-hour. And if you want to actually get your hands on a cluster of these chips, you’re looking at 12 to 18 months of waiting.
The price spike represents a sharp reversal from the trend that defined much of 2025, when GPU rental prices dropped more than 60% from their 2023-2024 peaks.
What’s driving the squeeze #
Inference workloads and multi-agent AI systems have pushed compute requirements well beyond what the existing supply base can handle. On-demand H100 capacity is effectively sold out across Neoclouds and hyperscalers alike.
Even modest deployments have become hard to secure. Clusters as small as 8 nodes, or 64 GPUs, are increasingly difficult to procure on short notice.
February 2026 saw particularly aggressive price moves, with 15-20% increases recorded in a single month as providers adjusted to the new demand reality.
The broader market for on-demand H100 rentals now spans a wide range, from $2.19 to over $4 per GPU-hour depending on the provider and whether the capacity is reserved in advance or purchased on the spot.
Blackwell backlog deepens the problem #
Lead times for B200 and GB200 deployments have stretched into mid-2026, with the majority of available 2026 capacity already spoken for.
Some operators have locked in their current H100 contracts at legacy rates, with renewals happening at previously agreed-upon pricing. A few have even extended their commitments through 2028. That kind of four-year contract length was nearly unheard of in cloud GPU markets just 18 months ago.
A market that forgot how to cool off #
Prices spiked during the initial generative AI boom in 2023-2024, then came a 60%-plus price decline in 2025, driven by aggressive capacity buildouts from Neocloud specialists like CoreWeave, Lambda, and others, alongside expanded offerings from the major hyperscalers.
By mid-2026, pricing showed some early signs of stabilization, though at levels well above where they sat six months prior.
What this means for the AI ecosystem #
For GPU cloud providers, higher utilization rates and rising prices translate to improved revenue per rack. Companies that locked in long-term supply agreements with Nvidia at favorable terms are sitting on what amounts to a spread trade: cheap wholesale, expensive retail. Hyperscalers like AWS, Google Cloud, and Azure have the balance sheets to absorb temporary margin compression, but smaller Neoclouds may find themselves in a stronger negotiating position than expected. When every GPU is spoken for, the provider with available capacity has leverage regardless of brand name.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our