The memory wall is seen as the biggest hurdle for rapidly scaling AI deployments, and moving memory closer to compute is seen as the optimal solution to the problem.
Astera Labs’ update to its Leo Smart Memory Controller family aims to do exactly that. The company’s new Leo X-Series fabric-attached memory controller, along with the next-generation Leo 2 E-Series and P-Series controllers, is designed to optimize memory performance as LLMs evolve into multi-turn agentic systems.
In a briefing with EE Times, Thad Omura, senior VP of Astera Labs’ compute connectivity group, said memory is the central infrastructure challenge as it races to keep up with the advances in GPU clustering during the early AI boom. “All the bottlenecks around memory are back and driving new architectures, and we’re trying to deal with the memory constraints,” he said.
The Leo 2 E-Series and P-Series continue Astera Labs’ CXL-based strategy for CPU-attached memory expansion and pooling. The devices support both DDR5 and reused DDR4 memory, offering up to 4 TB of DDR5 capacity per controller and up to 768 GB of repurposed DDR4. The company has also added memory testing, telemetry, predictive failure analysis, and automated repair capabilities to help hyperscalers redeploy aging DIMMs.
View All In benchmark testing using four Nvidia H200 GPUs running the Qwen2.5-32B model with multi-turn agentic workloads, Astera Labs reported a 62% reduction in time-to-first-token (TTFT) and a 22% increase in tokens per second compared with conventional CPU-DRAM-based KV cache off.
The memory wall is not the only pressure faced by AI adopters. In the same briefing, Ahmad Danesh, associate VP of product management, said DDR5 pricing has risen sharply as memory manufacturers prioritize high-bandwidth memory production. “There’s this big supply crunch and wafer supply crunch because of AI.”
Dynamic pooling and sharing features in the Leo 2 family are designed to reduce stranded memory and allow capacity to be allocated across servers as needed. “The expanded Leo family turns idle, stranded capacity into memory an AI agent can actually use, on hardware operators already own,” Omura said. “That stranded DRAM is the industry’s next scale-up resource.”
The Leo X-Series was developed specifically for inference workloads in which key-value (KV) cache data can overwhelm GPU high-bandwidth memory. Rather than routing data through CPU memory, the controller works in conjunction with Astera Labs’ Scorpio fabric switches to create a dedicated memory tier attached directly to the PCIe fabric.
Scorpio’s role is to optimize connectivity to AI accelerators and improve performance per watt, an increasingly important metric for data center operators. Astera Labs announced its latest Scorpio fabric switch earlier this year. In a briefing, Omura said fabric is the difference between peak and wasted compute.
He said the Scorpio X-Series 320-Lane Smart Fabric Switch reflects the company’s core mission of delivering purpose-built connectivity for rack-scale AI.
“Everything is measured in gigawatts,” Omura added, with scalability and resilience being equally important metrics, especially as inference grows at unprecedented scale.
Danesh said the latest Scorpio switch was built with large AI clusters in mind. It reduces latency by having a single hop for GPU communication.
To deliver that improved performance with lower power, Astera Labs has integrated dedicated hardware acceleration engines inside the chip, Danesh said. “Those engines are built to maximize the token economics further and get better performance per watt.”
He said Astera Labs’ hardware-accelerated and in-network compute engines can boost collective operations by up to 2×. For example, the in-network compute engine offloads some calculations off the accelerator. “We can calculate inside of our silicon,” Danesh said. “You’re spending less time sending the data and more time on the compute, which ultimately increases the GPU utilization.”
In addition to dedicated hardware inside Astera Labs’ silicon, the company’s Cosmos software adds another layer of intelligence. “The scale at which things are being deployed, it’s that software that really provides a lot of value,” Danesh said.
Danesh said the combination of hardware and software reflects a broader shift in AI infrastructure design; interconnects are no longer passive plumbing, but active system components that shape GPU utilization, quality of service, resilience, and the economic return on data center power budgets.
CXL has become an essential tool unlocking memory capacity to meet AI workload demands. Astera Labs released its Leo CXL Memory Accelerator Platform five years ago to enable CPUs to access and manage CXL-attached DRAM and persistent memory, making the use of centralized memory resources more efficient and allowing that access to scale up without slowing down performance.
South Korean startup Xcena also sees optimizing data movement as a critical path to address memory bottlenecks with its recently unveiled CXL Type 3 device that combines up to 2 TB of DDR5, SSD-backed InfiniteMemory capacity, and more than 1,000 custom RISC-V cores to push compute into memory.
Synopsys’s latest CXL offerings also address constrained memory capacity and bandwidth, while aiming to make the protocol easier to adopt. Marvell, meanwhile, released a suite of memory infrastructure products that aim to help hyperscale and cloud customers move memory closer to the point of computation, reduce data bottlenecks, and improve token efficiency for inference workloads.
Also read:
[Xcena Cuts Data Movement to Address Memory Bottlenecks](https://www.eetimes.com/xcena-cuts-data-movement-to-address-memory-bottlenecks/)
[Synopsys Updates CXL IP Portfolio for AI-Era Infrastructure](https://www.eetimes.com/synopsys-updates-cxl-ip-portfolio-for-ai-era-infrastructure/)
[Meta Recycles DDR4 Memory: Can Others Follow?](https://www.eetimes.com/meta-cuts-server-count-25-by-reusing-old-memory-can-anyone-else-do-it/)