# VMware Intros Private AI Cloud, AI Factory As Workloads Shift To On-Prem

> Source: <https://www.nextplatform.com/cloud/2026/09/01/vmware-intros-private-ai-cloud-ai-factory-as-workloads-shift-to-on-prem/5293559>
> Published: 2026-09-01 04:43:05+00:00

# VMware Intros Private AI Cloud, AI Factory As Workloads Shift To On-Prem

Ongoing cost, sovereignty, and security issues that arise as enterprises move from experimentation with AI to large-scale production deployments are continuing to fuel the shift of AI workloads away from public cloud environments to [on-premises private clouds and AI factories](https://www.nextplatform.com/enterprise/2026/06/22/hpe-rides-the-agentic-ai-wave-back-into-the-datacenter/5259579), which are becoming the dominant datacenter model for many organizations. Research firm Omdia is predicting global datacenter investments will climb to almost $1.6 trillion by 2030, and that this year, tech enterprises will spend [more than $600 billion](https://omdia.tech.informa.com/pr/2026/may/ai-factory-market-enters-its-industrialization-era-5-dynamics-redefining-ai-infra-in-2026) in AI infrastructure.

It’s a significant jump from traditional datacenters to AI factories, which Omdia defines as infrastructure designed to produce intelligence, with token production at its center and “characterized by ultra-high capital intensity, strong geopolitical attributes, and complex engineering barriers.”

Nvidia coined the term AI factory, which is just supercomputers running GenAI inference that crank out tokens to make money directly or to help get it in some fashion. OEMs like Dell, Hewlett Packard Enterprise, and others offer their own AI factories, with many based largely on Nvidia GPUs and other offerings.

At the VMware Explore 2026 event this week in Las Vegas, VMware and its owner, [Broadcom](https://www.nextplatform.com/connect/2026/03/05/broadcom-may-become-the-biggest-counterbalance-to-nvidia/4093625), are more tightly embracing the shift to on-premises AI production environments, building on its VMware Cloud Foundation 9 (VCF 9), [introduced two years ago](https://www.nextplatform.com/control/2024/08/27/vmware-wants-to-redefine-private-cloud-with-vcf-9/1633908), with the introduction of its VMware Private AI Cloud and – at its foundation – the VMware AI Factory model-as-a-service offering.

In VCF 9, VMware introduced a number of features to address such challenges as cost and management of AI workloads, including NVMe memory tiering that company executives say can reduce per-host costs by as much as 42 percent and VMware AI Assistant, an interface for resolving complex issues arising from such areas as CPUs, memory, and hypervisors.

The latest moves – which includes enhancements to Tanzu and more capabilities around the critical security issue in AI – will help VMware users to more easily move their AI workloads into their own environments, according to Prashanth Shenoy, vice president of marketing for Broadcom’s VMware Cloud Foundation Division.

“We've seen as organizations are moving from pilot deployments of AI applications and workloads to more production deployments of AI done at scale, they are facing major challenges around cost and security when it comes to data privacy, resiliency, availability of the infrastructure,” Shenoy told journalists during a pre-Explore media briefing. “That is causing a lot of organizations to repatriate workloads and run some of the production AI workloads back to an on-premises and private cloud environment.”

He pointed to the results in [VMware’s June survey](https://www.vmware.com/docs/private-cloud-outlook-2026) of 1,800 IT decision-makers, with 56 percent saying that their organizations are running or planning to run production inference in a private cloud. Sixty-two percent are concerned about costs and 51 percent said they’re repatriating AI workloads due to security concerns. The survey also found that use of public clouds for such workloads fell 15 percent year-over-year to 41 percent.

“They are preferring private cloud to run their production workloads – be it inferencing, fine-tuning, rack kind of AI use cases – in a private cloud on-premises because of the cost and security environment,” he said, adding that VMware Private AI Cloud “helps address the cost efficiency issues that we can bring to our customers, the security in the frontier AI model-driven security landscape, as well as customers moving to an agentic way of building their AI applications as well as workflows.”

VCF 9 already supports CPUs, GPUs, and other accelerators from multiple vendors and server hardware from OEMs and ODMs. Now comes Private AI Cloud, with VMware AI Factory at its foundation. VMware AI Factory includes VCF AI ReadyNodes from such hardware makers as [Dell Technologies](https://www.nextplatform.com/compute/2026/06/01/dell-makes-the-profits-up-in-volume-for-booming-ai-servers/5249707), [Cisco Systems](https://www.nextplatform.com/ai/2026/06/03/cisco-preps-for-a-world-of-ai-agent-coworkers-frontier-model-threats/5250406), Lenovo, and [Supermicro](https://www.nextplatform.com/compute/2026/08/13/the-genai-boom-will-lift-supermicro-but-it-will-lift-others-too/5287161), among others. In addition, it leverages [AMD technologies](https://www.nextplatform.com/compute/2026/08/05/amd-catches-the-agentic-ai-wave-and-will-ride-it-up-masterfully/5283468), including its Instinct MI350 Series GPUs and open ROCm, an open software platform for GPU accelerators, AI, and HPC on AMD hardware.

The AI factories have been tested by VMware and parent company [Broadcom](https://www.nextplatform.com/compute/2026/06/09/ai-chip-shepherds-broadcom-and-marvell-have-skinned-the-golden-fleece/5253011), Shenoy said.

“This is very critical because we have seen more and more customers wanting to run AI where their data lives, not the other way around,” he said. “They want to bring data to the model and not model to the data. But building this entire AI infrastructure from the metal to the module is a very manual process, where they have to stand up the infrastructure, put together the server with the right GPU size and form factor, bring the networking, the storage requirements, the Kubernetes for container workloads, the AI software stack, and different kind of models, be it SLMs [small language models], LLMs, open source, and open-weight models.”

Instead, VMware through the AI factories delivers an entire integrated package, all orchestrated via MetalSoft’s platform, which offers a unified operation of the underlying hardware and software.

VMware, through its VCF 9, also is expanding to more than 150 the [open and commercial AI models](https://www.nextplatform.com/ai/2026/08/13/the-war-between-open-source-open-weight-and-closed-ai-models/5287504) that can run on AI Factory can run, adding support for [Nvidia’s Nemotron 3](https://www.nextplatform.com/ai/2026/08/11/nvidia-drives-bang-for-the-buck-with-new-genai-model-and-router/5286365), Google DeepMind’s [Gemma 4](https://www.nextplatform.com/ai/2026/08/11/nvidia-drives-bang-for-the-buck-with-new-genai-model-and-router/5286365), Alibaba’s Qwen 3.7-Max, NEC’s cotomi, and GLM 5.2 from Z.ai.

VMware is introducing new AI services, including the ability to share AI models between tenants and lines of business via isolated nameplates, ensuring security and eliminating the need to deploy the same models, AI Gateway for ensuring unified model governance between cloud and on-premises environments, and secure AI sandboxes, virtualized container spaces for isolating agent code execution. Highly secure sandboxes are a focus in the agentic AI field after instances of OpenAI and Anthropic agents broke out of sandboxes to compromise third-party IT environments.

Other security controls include TrueSource by Broadcom for verifying open source AI, enhancements to vDefend to include zero-trust security for agentic AI workloads, and agentic threat defense via VMware’s Avi Load Balancer, another tool for ensuring that agents don’t access unauthorized tools, run zero-day attacks, or exfiltrate sensitive data without guardrails.

Similarly, VMware enhanced Tanzu, its agent platform for Private AI Cloud, with stronger security features that include hardened agent sandboxes with a “deny-by-default” containment model. It isolates credentials to prevent prompt injection attacks and network access that isn’t authorized by users. AI-ready data foundations processes both structured and unstructured data on-site to help improve agents’ accuracy, lower the cost of tokens, and reduce hallucinations.

A ready-to-go harness is aimed at giving developers pre-approved skills for agents, workflow buildpacks, human-in-the-loop controls, and integrated memory services. There also is marketplace where developers and agents can security connected to vetted AI models and tools, while an AI gateway is available for monitoring and logging all actions an agent takes for compliance and auditing purposes.

Shenoy said such features are critical “because all of these autonomous AI agents generate and execute code dynamically, and that creates security risk without proper governance as to, ‘What can these agents access from a tool, from a data, which agents can talk to which other agent.’ This framework will sandbox the agents within isolated secure containers and regulate which tools and agents they can communicate with.”
