cd /news/ai-infrastructure/ai-inference-gets-a-new-tier-as-cont… · home topics ai-infrastructure article
[ARTICLE · art-110755] src=siliconangle.com ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

AI inference gets a new tier as context windows grow

Solidigm Inc., Vast Data Inc., and Super Micro Computer Inc. are addressing new AI storage infrastructure demands as agentic AI expands context windows and KV caches, with Solidigm's Scott Shadley noting that SSDs have found a new home in a '3.5 tier' that enables faster data access during inference. Vast Data's Anat Heilper highlighted that high KV cache hit rates can save on expensive GPU compute and reduce latency, while Supermicro's Context Memory eXension (CMX) proposal targets large AI clusters.

read5 min views2 publishedAug 25, 2026
AI inference gets a new tier as context windows grow
Image: Siliconangle (auto-discovered)

AI inference gets a new tier as context windows grow

AI storage infrastructure is becoming a more consequential planning issue as organizations move from model training toward agentic AI. As agents reason, act and reassess, they build longer contexts and generate more data that they must access quickly during inference.

Agentic AI is also changing the shape of the data problem. Interactions are growing longer and producing more information. At the same time, the data’s size, importance and movement through the infrastructure can all affect how quickly an application responds, according to Scott Shadley (pictured, left), director of technology planning at Solidigm Inc.

“One of the beautiful things that’s happened in this agentic AI, or even just the AI era, is [that] people are starting to pay attention to storage,” he said. “What’s unique about this particular era is it’s no longer one- or two-dimensional. We have data magnitude and growth in size, importance and all of the other volumetric aspects of that. But at the end of the day, it comes down to that bit of data and how fast that bit of data moves from point A to point B.”

Shadley, along with Anat Heilper (center), director of AI architecture at Vast Data Inc., and Ben Lee (right), director of solution management at Super Micro Computer Inc., spoke with theCUBE Research’s Rob Strechay during the Supermicro Open Storage Summit interview series. They discussed how the companies’ respective technologies fit together as expanding context windows and KV caches create new storage and memory demands for agentic AI.

AI storage infrastructure brings KV cache closer to compute

As context windows expand, graphics processing unit memory alone can’t hold everything an agentic workload needs during inference. AI storage infrastructure must therefore provide additional tiers that balance proximity, capacity and speed, with each layer handling a different part of the data load, according to Shadley.

“As you think through that architecture — and you need to put that context somewhere — that context can start living in what used to be a no-no zone,” he said. “So [solid-state drives] have found a new home. One of the unique things about this 3.5 tier that we’re creating is that it could not exist until we had things like [Non-Volatile Memory Express] SSDs.”

That middle tier is one part of a larger architecture. Solidigm supplies SSDs that provide fast access to cached data, Supermicro integrates them into rack-scale systems and Vast Data’s AI Operating System provides network storage and the data services needed to use, manage and protect that data alongside broader AI workloads, Heilper noted.

“When we talk about [key-value] cache, which is a very significant optimization that can be done in AI inferencing … in essence, it’s the ability to replace compute with storage,” she said. “This is very significant because we all know the GPU is very, very expensive. When you have very high KV cache hit rates, we both save on compute and reduce the latency significantly.”

AI infrastructure needs room to evolve

Supermicro’s Context Memory eXtension, or CMX, proposal targets organizations with large AI clusters and substantial data demands; other deployments may require different combinations of memory, local SSDs and network storage. The company’s broader value lies in composing those building blocks around each customer’s workload, rather than treating one architecture as a universal answer, according to Lee.

“We believe solving the problem will take the whole rack because you cannot just buy more GPUs with more [high-bandwidth memory] … it’s very expensive,” he said. “All the KV cache will naturally overflow from the GPU, HBM, to the system memory, to the local SSD and to the network storage. But there’s a new industrial definition to try and fill the gap, and they call it G3.5, which is the CMX solution. This is a very AI-native KV cache tier that can fulfill the demand.”

Vast Data has been testing KV cache offload with partners across different software environments. Those experiments aim to show how the architecture behaves when cached context is integrated into production inference workloads, according to Heilper.

“With Nvidia Dynamo, we’ve shown that we managed to get 20 times faster time-to-first-token, which means that the latency that you perceive as a user is significantly faster,” she said. “Also, [we’ve] seen 90% savings in GPU time.”

Those results depend on an AI storage infrastructure that can match storage performance and capacity to the workload. Solidigm’s D7-PS1010 performance-oriented SSD and D5-P5336 capacity drive address different points in the hierarchy, according to Shadley. Supermicro integrates those components into systems that can be configured around customer requirements.

“This is not the only definition of a hierarchy stack,” Shadley said. “It’s the current primary that everybody leverages as the gold standard, but it’s continuing to evolve, and it’s unique. Being very proactive with your customer or your supplier to better understand what they know about what you need is no longer transactional. Those value-level transaction conversations are now what are going to drive the future of us deploying these types of architectures.”

Stay tuned for the complete video, part of SiliconANGLE and theCUBE’s coverage of the Supermicro Open Storage Summit interview series.

( Disclosure: TheCUBE is a paid media partner for the Supermicro Open Storage Summit interview series. Neither Supermicro, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)*

Photo: SiliconANGLE

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

About SiliconANGLE Media

SiliconANGLE,

theCUBE Network,

theCUBE Research,

CUBE365,

theCUBE AIand theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @solidigm inc. 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-inference-gets-a-…] indexed:0 read:5min 2026-08-25 ·