# VAST Data and AMD Claim 9x Faster Time-to-First-Token With KV Cache Offload on Instinct

> Source: <https://www.storagereview.com/news/vast-data-and-amd-claim-9x-faster-time-to-first-token-with-kv-cache-offload-on-instinct>
> Published: 2026-07-27 15:18:22+00:00

VAST Data has expanded its collaboration with AMD around AI infrastructure for cloud providers and enterprises deploying training, inference, retrieval-augmented generation, and agentic AI services. The effort combines the VAST AI Operating System with [6th Gen AMD EPYC processors](https://www.storagereview.com/news/amd-6th-gen-epyc-venice-256-cores-1-6tb-s-and-the-first-pcie-gen-6-server-cpu), AMD Instinct GPUs, AMD networking hardware, and ROCm software.

The collaboration reflects a broader infrastructure shift from training-focused AI clusters toward environments that must also support persistent context, high-concurrency inference, and multi-turn agent workflows. These deployments place increased emphasis on data movement, KV cache management, GPU utilization, and the ability to deliver low-latency access to large model and context datasets.

VAST positions its [Disaggregated Shared Everything, or DASE](http://VAST Data Introduces DPU-Native Inference Architecture for Shared KV Cache and Long-Lived Agentic AI), architecture as the shared data layer for these environments. The platform combines file and object storage, databases, event streaming, and data services under a unified global namespace. It also provides multi-tenancy and workload isolation capabilities for AI clouds operating concurrent customer and application workloads.

### EPYC 9006 Platforms for VAST CBox and EBox Systems

VAST has selected 6th Gen AMD EPYC processors, formerly codenamed Venice, for its next-generation CBox and EBox platforms. The company plans to use the processors in its sixth-generation CBox and third-generation EBox systems, which underpin the VAST AI Operating System.

[AMD EPYC 9006](https://www.storagereview.com/news/amd-6th-gen-epyc-venice-256-cores-1-6tb-s-and-the-first-pcie-gen-6-server-cpu) support introduces PCIe Gen6 connectivity to the VAST hardware platform. VAST states that the new interface doubles generational I/O bandwidth, improving file and object storage throughput while reducing latency for data services such as databases, data warehouses, and event-streaming workloads delivered through VAST DataBase and DataEngine.

For AI infrastructure, higher I/O bandwidth can help reduce bottlenecks between compute, network, and NVMe storage resources. This is particularly relevant for model loading, retrieval pipelines, checkpoint access, and externalized KV cache workflows where GPU memory capacity alone is insufficient to retain active context.

### Reference Architecture Combines Helios, VAST, and DriveNets

VAST, AMD, and [DriveNets](https://www.storagereview.com/news/whitefibers-project-redwood-links-two-h200-clusters-into-one-111-2-tbps-supercluster) are also developing an AI infrastructure reference architecture built around AMD Helios rack-scale AI infrastructure, the VAST AI Operating System, and DriveNets AI Fabric networking.

The reference architecture is intended to provide deployment guidance for model training, inference, reinforcement learning, and KV cache workloads. It includes sizing and availability considerations for organizations building AI factories that require shared data infrastructure and high-performance networking alongside GPU compute.

VAST also cited expanded collaboration with inference software providers TensorMesh and EmbeddedLLM. The ecosystem effort is focused on production inference architectures for agentic AI applications, although specific product integrations and availability details were not disclosed.

### KV Cache Offload Targets High-Concurrency Inference

A central component of the announcement is expanded [KV cache](http://The Token-Efficient Path for Long-Context Inference: KV Cache Offload to Flash) support using [AMD Instinct GPU](http://AMD MI455X and Helios: 432GB HBM4, 72-GPU Racks, and a Real Answer to Vera Rubin)s, AMD Infinity Context, ROCm software, and the VAST AI Operating System.

KV cache data stores intermediate attention-state information generated during inference. Retaining this data improves performance for multi-turn interactions and long-context workloads, but large cache footprints can consume significant GPU memory. Externalizing or tiering KV cache to a high-performance storage platform can free GPU memory for active workloads while retaining context for subsequent inference requests.

VAST reported early testing with an [AMD Instinct MI355X GPU](https://www.storagereview.com/review/supermicro-jumpstart-review-h14-with-amd-instinct-mi350x) that showed a 9x improvement in time to first token and 9.7x higher token throughput when using VAST for KV cache offload in high-concurrency agentic AI workloads. The company noted that these results depend on the hardware baseline, workload characteristics, and storage configuration.

The integration also applies VAST lifecycle policies to KV cache data. This capability is intended to automatically expire and delete cached information, which is relevant when the inference context contains sensitive, personal, or regulated data.

### Pensando Pollara 400 Connects GPUs to Shared Storage

The architecture uses the [AMD Pensando Pollara 400 AI NIC](https://www.storagereview.com/news/supermicro-h15-servers-pair-6th-gen-epyc-with-mi350p-gpu-systems-and-the-helios-rack) to connect AMD Instinct GPUs with the VAST storage cluster. According to VAST, the NIC supports GPU-to-storage data movement through NFS over TCP and RDMA, enabling access to NVMe SSD-based VAST clusters for KV cache and broader inference data requirements.

The resulting design is intended to support AI environments that need to scale context management independently of GPU memory. Rather than treating storage as a separate persistence layer, VAST is positioning the platform as a data and execution layer that can manage models, databases, streaming data, file data, and AI context across distributed infrastructure.

For AI cloud providers, the approach targets higher GPU utilization and improved operational efficiency as deployments move beyond GPU rental and batch training into persistent inference and agentic AI services.
