# Liqid Pools 30 AMD MI350P GPUs in One Server: 4.3TB of HBM3E and 69 PFLOPS for AI Inference

> Source: <https://www.storagereview.com/news/liqid-pools-30-amd-mi350p-gpus-in-one-server-4-3tb-of-hbm3e-and-69-pflops-for-ai-inference>
> Published: 2026-08-12 18:40:29+00:00

In late July, Liqid unveiled a scale-up AI platform built around AMD Instinct MI350P PCIe GPUs. The UltraStack 30 pairs Liqid’s composable GPU pooling technology with AMD accelerators to address enterprise and cloud AI inference deployments that need to scale beyond the accelerator count and memory capacity of conventional servers.

## UltraStack 30: Pooling 30 AMD MI350P GPUs

The proposed architecture enables up to 30 AMD Instinct MI350P GPUs to be pooled within a single server environment. Liqid states that this can provide 4.3TB of aggregate HBM3E capacity for large models, retrieval-augmented generation deployments, and long-context inference workloads. The platform is intended to allow GPU resources to be allocated dynamically, rather than fixed to individual servers, with native Kubernetes support for parallel model deployments.

The collaboration targets a growing infrastructure challenge in AI inference: maximizing useful accelerator utilization while controlling the cost of generated tokens. Liqid claims its pooled GPU approach can deliver up to 3.7x higher token throughput, 2.1x more tokens per dollar, and 1.8x better tokens per watt compared to traditional server configurations. The company also targets up to a 65% reduction in deployment cost and a 50% reduction in power consumption for large-model deployments, figures it says are based on internal projections that vary by configuration and workload.

“Liqid has established itself as the leader in GPU pooling and scaling. The AMD Instinct MI350P Series gives us the ideal PCIe-based GPU to build the next generation of AI infrastructure, allowing us to deliver solutions tuned for enterprise AI inference where utilization and cost per token decide the economics,” said Rick Hegberg, CEO of Liqid.

Liqid UltraStack 30: AMD MI350P |
|
Specification |
Configuration |
| Server | AMD EPYC™ 9005 Series CPU, Dual-Socket |
| GPUs | 30× AMD Instinct™ MI350P PCIe |
| AI Performance | 69 PFLOPS (FP8) |
| GPU HBM Memory | 4.3 TB |
| Total System Power | ~22 kW |

Key Metric |
Value |
| Performance per kW | ~3.14 PFLOPS/kW |
| HBM per kW | ~195 GB/kW |
| HBM per PFLOP | ~62 GB/PFLOP |

## Inside the MI350P

AMD positions the [MI350P](https://www.storagereview.com/news/amd-instinct-mi350p-enterprise-pcie-ai-inference-returns-to-standard-servers) as a PCIe-based accelerator designed to deliver AI inference performance without requiring specialized cooling or a broader data center redesign. AMD’s spec sheet lists the card at 144GB of HBM3E with 4TB/s of memory bandwidth, 2.3 PFLOPS of dense FP8 compute from its CDNA 4 architecture, and a 600W typical board power (450W configurable) in a passively cooled, double-slot PCIe 5.0 form factor. Thirty of those cards is exactly where the UltraStack’s 4.3TB and 69 PFLOPS aggregate figures come from. Combined with Liqid’s fabric, the GPUs can be added and scaled on demand to support changing inference demand, potentially reducing stranded GPU capacity and enabling higher model density per system.

“The AMD Instinct MI350P PCIe GPU delivers exceptional AI inference performance without requiring specialized infrastructure. Together with Liqid’s GPU pooling solutions, customers can scale GPU resources on demand to achieve higher utilization, lower infrastructure costs, and industry-leading AI inference economics,” said Suresh Andani, corporate vice president, Compute and Enterprise AI Group at AMD.

## Who the Platform Targets

The solution is aimed at enterprise AI teams, NeoCloud providers, AI service providers, and HPC and research organizations. Identified use cases include private, on-premises enterprise assistants for functions such as sales, HR, marketing, legal, and coding; RAG and long-context inference; scientific workloads including drug discovery and materials science; and mixed-model fleets that serve multiple model sizes from a shared accelerator pool.

## From GPU Pooling to CXL Memory Pooling

The [CXL](https://www.storagereview.com/news/synopsys-cxl-4-0-ip-hits-128-gt-s-claims-3-6x-kv-cache-offload-over-ssds) side of that story is no longer theoretical. In early August, Liqid launched the EX-5410C Memory Platform, which the company calls the industry’s first and only fully disaggregated, software-defined memory pooling solution. Built on CXL 2.0 and managed through Liqid Matrix software, the EX-5410C pools up to 40TB of DRAM per chassis, scales to a unified pool exceeding 160TB, and dynamically allocates memory across as many as 16 server nodes with what Liqid describes as zero stranded capacity.

Liqid claims up to 30x faster processing for graph analytics workloads and up to 7x more tokens per second for certain KV cache workloads on the platform, and an early testbed is already running at Pacific Northwest National Laboratory, available to DOE-funded researchers through the AMAIS initiative. For the UltraStack, the takeaway is directional: the same composability model Liqid applies to GPU capacity is extending to system memory, which would let future deployments allocate shared memory for [KV cache](https://www.storagereview.com/review/the-token-efficient-path-for-long-context-inference-kv-cache-offload-to-flash) and other memory-intensive AI workloads alongside pooled accelerators.

Further details on joint Liqid and AMD solutions are expected as the companies advance development and customer deployment activities.
