{"slug": "vast-uses-tiered-storage-to-ease-ai-agent-memory-demands", "title": "Vast uses tiered storage to ease AI agent memory demands", "summary": "Vast co-founder and CTO Alon Horev said a long-running AI agent session of half a million tokens can consume one-tenth to one-twentieth of a GPU's memory, and that Vast's tiered storage approach offloads KV cache from GPU memory to CPU memory and then to persistent media holding petabytes, orchestrated by Nvidia's Dynamo software. Horev, speaking with theCUBE Research's John Furrier and Dave Vellante at Fully Connected 2026, said moving sessions across a fleet of machines lets enterprises avoid repeat recalculation and gain more optimized scheduling as they deploy thousands of agents handling sensitive data.", "body_md": "### Vast uses tiered storage to ease AI agent memory demands\n\nAI agent memory is creating new demands on infrastructure as agents run longer sessions and spread across the enterprise. Retaining that context and making it available when needed puts pressure on memory capacity and data movement.\n\nThose demands extend beyond the context held during an individual interaction. Enterprise agents also need shared knowledge that persists across sessions, according to [Alon Horev](https://www.linkedin.com/in/alonhorev/), (pictured) co-founder and chief technology officer of Vast.\n\n“Memory for agents is a bit different,” Horev said. “First of all, there are multiple types of memory. There’s long-term memory where an agent can see past conversations and past interactions and look back and learn from its past experiences.”\n\nHorev spoke with theCUBE Research’s [John Furrier](https://www.linkedin.com/in/furrier) and [Dave Vellante](https://www.linkedin.com/in/dvellante/) at [Fully Connected 2026](https://www.thecube.net/events/coreweave/coreweave-fully-connected-2026), during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed AI agent memory, key value cache offloading and the shift toward data movement as the next bottleneck. *(* Disclosure below.)*\n\n### Why AI agent memory is moving beyond the GPU\n\nThe pressure shows up first in inference. Each long-running session holds its KV cache in graphics processing unit memory, and a session of half a million tokens can take up one-tenth to one-twentieth of a GPU’s memory, Horev explained.\n\n“It’s also possible the agent would stop talking to the [large language model] because it’s compiling code, it’s testing software, or, as a human, I want to have a cup of coffee,” he said. “What you see is that if you could stretch that memory wall and basically offload those sessions to storage, you can avoid that repeat recalculation.”\n\n[Vast’s approach](https://www.vastdata.com/blog/beyond-hbm-limits) uses memory in tiers. GPU memory is used first, then central processing unit memory on the same machine, then persistent media that can hold petabytes of KV cache, with Nvidia Corp.’s Dynamo software orchestrating the process, Horev noted.\n\n“You can move a session from one busy GPU to one less busy GPU and move KV cache either over the network or read it from Vast,” he said. “Once you look at inference as a distributed problem where you have the opportunity to use GPU memory, CPU memory and Vast across a fleet of machines, you have more optionality and you have more optimized scheduling.”\n\nThe stakes rise as companies deploy thousands of agents that handle sensitive data and act on customers’ behalf. Those enterprises need to record everything their agents do and retain it for a set period, which makes AI agent memory both a governance and performance asset, according to Horev. Vast has also [launched a confidential computing service](https://siliconangle.com/2026/09/22/vast-data-launches-confidential-computing-service-for-sensitive-workloads/) for sensitive workloads.\n\n“These conversations that the agent is doing, it’s also gold,” he said. “It’s the same information that it can use for fine-tuning or training or creating purpose-built models.”\n\nHere’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of [Fully Connected 2026](https://www.thecube.net/events/coreweave/coreweave-fully-connected-2026):\n\n*(* Disclosure: TheCUBE is a paid media partner for the Fully Connected event. Neither CoreWeave, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)*\n\n##### Photo: SiliconANGLE\n\n# A message from John Furrier, co-founder of SiliconANGLE:\n\nSupport our mission to keep content open and free by engaging with theCUBE community. **Join theCUBE’s Alumni Trust Network**, where technology leaders connect, share intelligence and create opportunities.\n\n- **15M+ viewers of theCUBE videos** , powering conversations across AI, cloud, cybersecurity and more\n- **11.4k+ theCUBE alumni** — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network\n\n### Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: [https://siliconangle.com/aws-marketplace/](https://siliconangle.com/aws-marketplace/)\n\n##### **About SiliconANGLE Media**\n\n[SiliconANGLE](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552),\n\n[theCUBE Network](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da),\n\n[theCUBE Research](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f),\n\n[CUBE365](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6),\n\n[theCUBE AI](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683)and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.\n\nFounded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.", "url": "https://wpnews.pro/news/vast-uses-tiered-storage-to-ease-ai-agent-memory-demands", "canonical_source": "https://siliconangle.com/2026/10/06/ai-agent-memory-belongs-storage-says-vast-data-cto-fullyconnected/", "published_at": "2026-10-06 17:02:26+00:00", "updated_at": "2026-10-06 17:18:48.297422+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "ai-chips", "large-language-models"], "entities": ["Vast", "Alon Horev", "Nvidia", "Dynamo", "theCUBE", "SiliconANGLE", "John Furrier", "Dave Vellante"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/vast-uses-tiered-storage-to-ease-ai-agent-memory-demands", "markdown": "https://wpnews.pro/news/vast-uses-tiered-storage-to-ease-ai-agent-memory-demands.md", "text": "https://wpnews.pro/news/vast-uses-tiered-storage-to-ease-ai-agent-memory-demands.txt", "jsonld": "https://wpnews.pro/news/vast-uses-tiered-storage-to-ease-ai-agent-memory-demands.jsonld"}}