Liqid unveiled its UltraStack 30 platform in late July, pooling up to 30 AMD Instinct MI350P GPUs in one AMD EPYC-based server. According to Liqid's datasheet, the configuration provides 4.3 TB of aggregate HBM3E memory and up to 69 PFLOPS of FP8 performance. Liqid projects higher token throughput and lower deployment costs than multi-server designs, although those figures vary by workload and configuration.
Liqid unveiled the UltraStack 30, a scale-up AI inference platform that pools up to 30 AMD Instinct MI350P PCIe GPUs within a single AMD EPYC-based server. Liqid's July datasheet lists up to 4.3 TB of aggregate HBM3E memory and 69 PFLOPS of FP8 performance for the configuration.
Converge Digest reports that the system combines dual-socket AMD EPYC 9005 Series processors with 30 MI350P accelerators through Liqid's software-defined GPU pooling architecture. StorageReview reports that the fully configured platform consumes approximately 22 kW and is intended for large-model inference, retrieval-augmented generation, long-context workloads, enterprise AI assistants, scientific computing, and mixed-model deployments.
A PCIe-based scale-up design
Liqid's Matrix software attaches and orchestrates GPUs through a high-speed PCIe fabric, according to the company's datasheet. The company describes native Kubernetes support for serving multiple models, while Converge Digest reports that the architecture is intended to dynamically allocate GPU resources rather than statically bind them to individual servers.
The system is notable for targeting a scale-up configuration beyond the 8-GPU ceiling common in conventional servers. Liqid argues in its datasheet that consolidating large models or multiple models into one server can reduce multi-node networking and software overhead. That is a vendor claim rather than an independently published benchmark result.
"Liqid has established itself as the leader in GPU pooling and scaling. The AMD Instinct MI350P Series gives us the ideal PCIe-based GPU to build the next generation of AI infrastructure, allowing us to deliver solutions tuned for enterprise AI inference where utilization and cost per token decide the economics," Rick Hegberg, Liqid's CEO, said in comments published by StorageReview.
Performance and cost claims
According to Liqid's datasheet, UltraStack 30 can provide up to:
- • 3.7x more tokens per second - • 2.1x more tokens per dollar - • 1.8x better tokens per watt - • 65% lower deployment cost and50% lower power consumption for large-model deployments
StorageReview characterizes these as internal projections that depend on workload and configuration. The retrieved sources do not provide benchmark methodology, model sizes, batch sizes, interconnect comparisons, or measurements against named multi-server systems. Those omissions make it difficult to compare the claimed token economics with systems using dedicated scale-up fabrics or alternative GPU pooling implementations.
For infrastructure teams, aggregate HBM capacity is likely to be the central specification. Large memory pools can reduce the need to partition model weights and KV caches across multiple nodes, although realized latency and throughput depend on model parallelism, PCIe topology, scheduler behavior, request mix, and software support. Companies evaluating comparable pooled-accelerator designs typically need workload-specific measurements for prefill, decode, concurrency, and tail latency rather than peak FP8 figures alone. Liqid's datasheet also describes the platform as CXL memory-pooling ready, framing it as a path toward shared terabyte-scale memory for KV cache and other memory-intensive workloads. The source material does not identify a shipping CXL memory configuration or provide performance results for that capability.
Key Points #
- 1UltraStack 30 aggregates 30 MI350P GPUs and 4.3 TB of HBM3E, concentrating large-model inference capacity in one server.
- 2Liqid's throughput, cost, and power figures are internal projections, so practitioners need model-specific benchmarks before comparing platform economics.
- 3Comparable pooled-GPU systems shift evaluation toward memory capacity, topology, scheduler behavior, and tail latency rather than peak FLOPS alone.
Scoring Rationale #
The platform presents a notable high-density AMD GPU configuration for enterprise inference, with unusually large aggregate HBM capacity in a single server. Its practical significance depends on independently reproducible performance, topology, and workload-level measurements that are not included in the retrieved material.
Sources #
Primary source and supporting public references used for this report.
[Primary sourceliqid.comLIQID UltraStack 30 - AMD MI350P](https://www.liqid.com/products/ultrastack-30-amd)
[Independent reportingconvergedigest.comLiqid Debuts 30-GPU AI Platform Powered by AMD Instinct ...](https://convergedigest.com/liqid-debuts-30-gpu-ai-platform-powered-by-amd-instinct-mi350p/)
View 1 more source #
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.