Startup Majestic Labs wants to solve problem causing KV caching scheme by getting rid of AI inferencing’s need for high-bandwidth memory.
Its Prometheus server idea is that costly GPUs should be replaced by its own Ignite AI Processing Units (AIUs); hybrid chips combining datacenter-class ARM cores with RISC-V vector and tensor engines, for LLM processing, in a single die, all operating in one shared memory space built from LPDDR6 DRAM chips. There would be from 8 to 128 TB of this disaggregated and unified memory accessed by from 1 to 12 AIUs sharing around 25.6 TB/s of bandwidth.
For comparison, an Nvidia DGX B300 AI server, with 8 x Blackwell GPUs, has 2.3 TB of total GPU memory (HBM3e) plus up to 4 TB of system memory (DDR5), and an up to 14.4 TB/s low-latency, interconnect bandwidth, meaning the Prometheus server has more than 50 times that GPU server’s HBM and 1.7x its bandwidth. Majestic’s founders say the GPU-HBM (high-bandwidth memory) combo twins very fast and expensive compute with limited capacity fast memory. It’s limited because it hooks up to the restricted GPU chip shoreline; 4 sides of an oblong, restricted cable length, and because of the difficulty of increasing HBM beyond 12 layers.
The GPU-HBM strategy is a dead-end and is fundamentally memory bound. It should be replaced with more cost-effective processors accessing a vastly increased memory pool accessed through scalable custom memory aggregation chiplets (MACs) across miniature copper cables up to ~1 meter in length. The MACs are placed near the board-mounted memory chips, and a single MAC fans out to many chips.
The memory space is contiguous and coherent. The AIUs inter-communicate via the memory and connect to the MACs in a mesh-like structure and the number of MACs, unlike the number of AIUs, is not specified.
Okay. So far so good, but … if there is up to 128 TB of memory and generally available LPDDR6 chips are 2 GB in capacity, that means a 128 TB Prometheus server has, if we use a 2 GB LPDDR6 die, up to 64,000 LPDDR6 dies. We don’t know how many such dies one of Majestic Labs MAC chiplets can support but we’re surely talking about a hundred or more MACs per server.
There can be up to four Prometheus servers in a standard 40U rack, with each server containing 12 AIUs. They will draw a total of 120 kW and use cold-plate, liquid cooling. Majestic Labs claims “One Majestic rack holds the fast memory capacity of 25 Nvidia NVL72 Vera Rubin racks at a fraction of the power. Organizations that could never justify hyperscaler infrastructure can now run any workload.” In fact, there can be up to “1000× more memory per processor.”
When memory is no longer a bottleneck you can “efficiently serve the most advanced and biggest frontier models with the longest contexts. Efficiently run Agentic AI, Reasoning, Graph Neural Nets, Tabular Nets, Video Generation and every new model” and ”support 100x more users per rack, massively reducing power consumption.”
Customers can run “multi-trillion parameter models [with] massive context windows, agentic systems and mixture-of-experts, all in a single system at a fraction of the power and cost.”
It reckons it could cost between 10 and 50 times less than an equivalent performance GPU server system when it ships next year, and use less electricity.
The Prometheus server is OCP-compliant, and will support PyTorch, vLLM, and OpenAI’s Triton inference frameworks so that current AI models using these frameworks can run straight away.
Bootnote
Majestic Labs was founded in 2023 in Tel Aviv by a trio of ex-Google and Meta chip-level design, build and ship people: CEO Ofer Shacham, President Sha Rabii and COO Masumi Reynders (COO). The company has around 40 employees in Tel Aviv and a Los Angeles site, and raised $100 million in an A-round in late 2025. It has licensed 3rd-party accelerator IP with a custom version being developed for the Ignite AIU accelerator core. Majestic says it has received significant orders from multiple customers. The target customers are large enterprises, neoclouds and hyperscalers.