{"slug": "majestic-labs-wants-to-solve-memory-bound-gpu-problem", "title": "Majestic Labs wants to solve memory-bound GPU problem", "summary": "Startup Majestic Labs claims its Prometheus server, built with Ignite AI Processing Units and LPDDR6 memory, can deliver up to 128 TB of unified memory and 25.6 TB/s bandwidth, more than 50 times the GPU memory of an Nvidia DGX B300 server at 1.7 times the bandwidth. The company says the GPU-HBM approach is a dead end and that its system will cost 10 to 50 times less than equivalent GPU servers when it ships next year, supporting multi-trillion parameter models with massive context windows.", "body_md": "# Majestic Labs wants to solve memory-bound GPU problem\n\nStartup [Majestic Labs](https://majestic-labs.ai/) wants to solve problem causing KV caching scheme by getting rid of AI inferencing’s need for high-bandwidth memory.\n\nIts Prometheus server idea is that costly GPUs should be replaced by its own Ignite AI Processing Units (AIUs); hybrid chips combining datacenter-class ARM cores with RISC-V vector and tensor engines, for LLM processing, in a single die, all operating in one shared memory space built from LPDDR6 DRAM chips. There would be from 8 to 128 TB of this disaggregated and unified memory accessed by from 1 to 12 AIUs sharing around 25.6 TB/s of bandwidth.\n\nFor comparison, an[ Nvidia DGX B300 ](https://www.nvidia.com/en-us/lp/ai/ai-factories-in-action-ebook/?ncid=pa-srch-goog-251433&_bt=803288465757&_bk=nvidia%20dgx&_bm=p&_bn=g&_bg=195676304635&gad_source=1&gad_campaignid=23716531603&gclid=EAIaIQobChMIycvl4P7olQMVyaJQBh2MFANSEAAYASAAEgJdovD_BwEhttps://majestic-labs.ai/)AI server, with 8 x Blackwell GPUs, has 2.3 TB of total GPU memory (HBM3e) plus up to 4 TB of system memory (DDR5), and an up to 14.4 TB/s low-latency, interconnect bandwidth, meaning the Prometheus server has more than 50 times that GPU server’s HBM and 1.7x its bandwidth.\n\nMajestic’s founders say the GPU-HBM (high-bandwidth memory) combo twins very fast and expensive compute with limited capacity fast memory. It’s limited because it hooks up to the restricted GPU chip shoreline; 4 sides of an oblong, restricted cable length, and because of the difficulty of increasing HBM beyond 12 layers.\n\nThe GPU-HBM strategy is a dead-end and is fundamentally memory bound. It should be replaced with more cost-effective processors accessing a vastly increased memory pool accessed through scalable custom memory aggregation chiplets (MACs) across miniature copper cables up to ~1 meter in length. The MACs are placed near the board-mounted memory chips, and a single MAC fans out to many chips.\n\nThe memory space is contiguous and coherent. The AIUs inter-communicate via the memory and connect to the MACs in a mesh-like structure and the number of MACs, unlike the number of AIUs, is not specified.\n\nOkay. So far so good, but … if there is up to 128 TB of memory and generally available LPDDR6 chips are 2 GB in capacity, that means a 128 TB Prometheus server has, if we use a 2 GB LPDDR6 die, up to 64,000 LPDDR6 dies. We don’t know how many such dies one of Majestic Labs MAC chiplets can support but we’re surely talking about a hundred or more MACs per server.\n\nThere can be up to four Prometheus servers in a standard 40U rack, with each server containing 12 AIUs. They will draw a total of 120 kW and use cold-plate, liquid cooling. Majestic Labs claims “One Majestic rack holds the fast memory capacity of 25 Nvidia NVL72 Vera Rubin racks at a fraction of the power. Organizations that could never justify hyperscaler infrastructure can now run any workload.” In fact, there can be up to “1000× more memory per processor.”\n\nWhen memory is no longer a bottleneck you can “efficiently serve the most advanced and biggest frontier models with the longest contexts. Efficiently run Agentic AI, Reasoning, Graph Neural Nets, Tabular Nets, Video Generation and every new model” and ”support 100x more users per rack, massively reducing power consumption.”\n\nCustomers can run “multi-trillion parameter models [with] massive context windows, agentic systems and mixture-of-experts, all in a single system at a fraction of the power and cost.”\n\nIt reckons it could cost between 10 and 50 times less than an equivalent performance GPU server system when it ships next year, and use less electricity.\n\nThe Prometheus server is [OCP-compliant](https://www.blocksandfiles.com/glossary/2022/05/04/ocp/1592446), and will support PyTorch, vLLM, and OpenAI’s Triton inference frameworks so that current AI models using these frameworks can run straight away.\n\n##### Bootnote\n\nMajestic Labs was founded in 2023 in Tel Aviv by a [trio](https://majestic-labs.ai/team) of ex-Google and Meta chip-level design, build and ship people: CEO Ofer Shacham, President Sha Rabii and COO Masumi Reynders (COO). The company has around 40 employees in Tel Aviv and a Los Angeles site, and raised $100 million in an A-round in late 2025. It has licensed 3rd-party accelerator IP with a custom version being developed for the Ignite AIU accelerator core. Majestic says it has received significant orders from multiple customers. The target customers are large enterprises, neoclouds and hyperscalers.", "url": "https://wpnews.pro/news/majestic-labs-wants-to-solve-memory-bound-gpu-problem", "canonical_source": "https://www.blocksandfiles.com/ai-ml/2026/07/23/majestic-labs-wants-to-solve-memory-bound-gpu-problem/5277163", "published_at": "2026-07-23 14:49:22+00:00", "updated_at": "2026-07-23 15:06:57.878116+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-chips", "ai-startups", "ai-products"], "entities": ["Majestic Labs", "Prometheus server", "Ignite AI Processing Units", "Nvidia DGX B300", "Blackwell GPUs", "LPDDR6", "RISC-V", "OpenAI Triton"], "alternates": {"html": "https://wpnews.pro/news/majestic-labs-wants-to-solve-memory-bound-gpu-problem", "markdown": "https://wpnews.pro/news/majestic-labs-wants-to-solve-memory-bound-gpu-problem.md", "text": "https://wpnews.pro/news/majestic-labs-wants-to-solve-memory-bound-gpu-problem.txt", "jsonld": "https://wpnews.pro/news/majestic-labs-wants-to-solve-memory-bound-gpu-problem.jsonld"}}