Via aboutamazon.com
Matt Garman says inference now accounts for two-thirds of AI compute demand, a trend that could reshape the economics of GPU-heavy operations across cloud and crypto mining sectors.
At AWS re:Invent 2025, AWS CEO Matt Garman laid out a simple but consequential data point: AI inference now accounts for roughly two-thirds of all AI compute demand. That’s up from about one-third in 2023.
In English: training is teaching a model to think. Inference is actually making it think. The industry spent billions teaching these models, and now the spending is shifting toward putting them to work.
Garman, who took over as AWS CEO in June 2024 after succeeding Adam Selipsky, described inference as “a new Lego” in computing. He’s saying inference is becoming a basic building block that enterprises will snap together to automate tasks, run AI agents, and fundamentally change how businesses operate.
He went further, predicting that 80-90% of enterprise AI value will eventually come from inference-powered agents. Not chatbots summarizing your emails, but autonomous systems actually completing tasks.
AWS is putting hardware behind this bet too. The company showcased its Trainium3 chips at re:Invent, purpose-built for inference workloads rather than the training-optimized GPUs that have dominated headlines and Nvidia’s earnings reports.
This shift matters enormously for the crypto mining sector, which has been in the middle of its own identity crisis. As Bitcoin mining margins tightened after the April 2024 halving, several major miners pivoted toward offering AI and high-performance computing services. Companies like Core Scientific, Hut 8, and others have been repurposing their GPU fleets and data center capacity to serve AI training workloads.
But if Garman is right, and the demand curve is bending hard toward inference, the hardware requirements change. Training demands massive parallel processing across clusters of top-tier GPUs running for weeks or months. Inference needs different optimization: lower latency, higher throughput per query, and often different chip architectures entirely.
For crypto-native investors, the signal here is about the sustainability of the “AI pivot” thesis that has propped up several mining stocks. Watch for which mining companies announce partnerships with cloud providers or invest in inference-optimized hardware. The ones still pitching generic GPU access for training may find themselves on the wrong side of a demand curve that’s already shifted. Decentralized compute networks like Render, Akash, and io.net have been positioning themselves as alternatives to centralized cloud providers for AI workloads. If inference demand is growing at the rate Garman describes, these protocols could see increased utilization, but only if they can offer the low-latency, high-reliability service that inference requires. Training is forgiving of network hiccups. Inference is not.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our