# Nutanix Adds Additional Capabilities to Optimally Run AI Workloads

> Source: <https://techstrong.ai/articles/nutanix-adds-additional-capabilities-to-optimally-run-ai-workloads/>
> Published: 2026-08-26 13:32:27+00:00

TL;DR — Key Takeaways

- Nutanix is adding MCP support to its AI gateway and broader cloud platform to simplify access to distributed data for AI workloads.
- Nutanix Enterprise AI 2.8 and the upcoming NKP 2.19 are designed to improve deployment, security and management of AI applications across virtualized and Kubernetes environments.
- New inference capabilities support tensor parallelism, batch inference and speculative decoding, potentially increasing token-generation rates by as much as 2.5 times.

Nutanix today added support for the Model Context Protocol (MCP) to an artificial intelligence (AI) gateway and its core integrated platform for deploying virtual machines and containers.

Additionally, Nutanix is adding a built-in catalog to its Nutanix Kubernetes Platform (NKP) that streamlines the deployment of AI applications that can be deployed on Nutanix Private Inference that has been enhanced to enable IT teams to better secure and fine tune the performance of AI models.

These extensions to the Nutanix Cloud Platform (NCP) will make it simpler for IT teams to deploy and manage AI workloads in self-hosted environments that span everything from on-premises IT platforms to private clouds, says Thomas Cornely, executive vice president of product management for Nutanix.

The overall goal is to ensure that all types of AI workloads, including agentic, can be reliably run in a production environment, he adds. AI agents, in particular, need to be deployed in IT environments that are able to dynamically scale up and down to meet requirements that are much less predictable than previous classes of workloads, notes Cornely. “Agentic workloads are putting a lot of pressure on things IT teams used to take for granted,” he says.

With the release of Nutanix Enterprise AI (NAI) 2.8 and, shortly, NKP 2.19, Nutanix, for example, provides access to an AI gateway based on open source Envoy proxy software that has been extended to make it easier via MCP to access distributed data at scale, adds Cornely.

Additional inference and fine tuning capabilities also enable scalable inference for multiple large language models (LLMs) running on graphics processing units (GPUs) via tensor parallelism. That capability enables batch inference and speculative decoding that can increase the rate at which tokens are generated by as much as 2.5 times using lightweight draft models.

NCP also prevents rogue AI models from being deployed using application programming interfaces (APIs) that use fine-grained Identity and Access Management (IAM), custom roles and model sharing in a way that prevents them from escaping a sandbox environment.

The issue IT teams are facing in the AI era is that they need to support multiple classes of workloads, says Cornely. NCP is specifically designed to enable IT teams to deploy AI applications on a Kubernetes cluster alongside legacy applications running on virtual machines, he notes. That approach enables organizations to centralize the management of workloads in a way that reduces the total cost of IT, says Cornely.

Alternatively, IT teams can also soon opt to deploy AI applications running on a NKP deployed on a bare-metal platform. Regardless of approach, NCP provides IT teams with the level of flexibility needed as the workloads being deployed continue to evolve, notes Cornely.

It’s not clear to what degree centralized IT teams are assuming more responsibility for deploying AI workloads in production environments, but there is little doubt that as the number of them increases, they will ultimately be managed much like any other workload. The challenge is that highly dynamic AI workloads running across a hybrid IT environment don’t behave much like any type of workload most IT teams have previously managed.
