TL;DR — Key Takeaways
- Managed AI inference spending overtook AI training spending in 2025, signaling a broader shift from experimentation to production deployments.
- Futurum Group projects managed inference will grow significantly faster than training through 2030 as enterprises deploy more AI workloads.
- Provider-managed cloud adoption increases as organizations move from AI pilots toward broader transformation, reflecting continued dependence on cloud providers for inference infrastructure.
The Futurum Group is reporting that at the close of 2025 the overall size of the managed artificial intelligence (AI) services for the first time became larger than the AI training market.
The managed inference market reached $23.1 billion at the end of 2025, while training accounted for $16.3 billion. Overall, the Futurum Group projects that by 2030 the managed inference base case is projected to reach $106.8 billion for a 36% compound annual growth rate (CAGR), compared to $42.2 billion spent on AI training, representing a 21% CAGR.
While it’s always been expected that investments in inference models would ultimately overtake training, the rate at which it has occurred is significantly faster than many initially expected, says Nick Patience, practice lead for AI Platforms at the Futurum Group. “The direction of travel has been obvious for a couple of years,” says Patience. “But we’ve seen the crossover complete in 2025. This indicates that enterprises moved from pilot to production quicker than the market expected.”
In fact, a Futurum Group survey of 820 enterprise AI decision-makers finds provider-managed cloud adoption rises from 57% at the experimentation stage to 73% at the transformation stage as AI workloads are deployed. That uptick suggests that, for now at least, IT organizations are depending mainly on cloud service providers to deploy their AI inference models.
There are, of course, plenty of IT organizations that have opted to self-host AI inference models, but regardless of approach the volume of AI workloads running in production environments has substantially increased. As those workloads increase, however, organizations are faced with the challenge of determining how to best manage them. While in many cases the AI inference models have been deployed by a cloud service provider, they are also generally co-managed by either a dedicated team or a centralized IT team that as of late is playing a larger role in the deployment of AI inference models. The primary challenge is that the amount of internal AI expertise most organizations have today is limited, so there is going to be a natural tendency to rely more on managed service providers (MSPs) to deploy AI models.
Regardless of how AI inference models are managed, there will soon be a lot more of them not just running in the cloud but also at the network edge and eventually on handheld devices. The challenge then becomes determining where best to deploy them based on the latency requirements of the application. It’s also probable that many inference models that are initially deployed in the cloud will either be moved closer to the network edge or be federated with other AI inference models that might have been distilled from the same core foundational model.
It’s hard to imagine at this point any new workload being deployed that to one degree or another isn’t going to be dependent on an AI inference model. The challenge and the opportunity now is to make sure the right AI inference model is running in the right place at the right time.