{"slug": "ai-memory-prices-and-the-engineering-response", "title": "AI, Memory Prices, and the Engineering Response", "summary": "Google said at Semicon Taiwan that it is redesigning its systems to cope with memory prices that now account for more than 75% of a server bill of materials, according to Google. The company's new chips hold more data close, share memory across thousands of processors, and compress AI working memory sixfold, creating substantially more effective memory capacity without waiting for new fabs. Compressing KV caches, which consume 15–40% of memory in long-context inference, could reduce total memory requirements by 12–33% and support 15–50% more inference work.", "body_md": "## What Happened\n\nThe most visible response to **soaring memory prices**—now perhaps more than **75% of a server bill of materials (** according to Google)—is higher semiconductor capex: new fabs, more HBM packaging, and an approaching flood of [Chinese DRAM capacity](https://paulkedrosky.com/memory-scarcity-is-creating-its-own-exit-part-ii-cxmt-ymtc-etc/).\n\nAs ever, **too much analysis is static**, naively **extrapolating** from the present and ignoring both **engineering and supply responses**. While the latter will come, the former is faster.\n\nGoogle described at Semicon Taiwan this week how it is** redesigning its systems** to get more from scarce memory. Its new chips **hold more data close** at hand, **share memory across thousands of processors**, and **compress the working memory used by AI models sixfold**. **Older, cheaper DRAM** can also take over less demanding jobs. Combined, these changes create **substantially more effective memory capacity** without waiting for new fabs.\n\nThe upshot: These changes create **a step change in more effective memory capacity**, with no costly cleanroom required (even if that's coming). And Google isn't the only company taking these steps.\n\n### What It Means\n\n**The changes are multiplicative**- If\n**KV caches consume 15–40% of memory** in a long-context inference workload.**Compressing them sixfold** reduces total memory requirements by roughly**12–33%**. - The same hardware could therefore support about\n**15–50% more inference** work. - The largest gains appear in\n**long-context and agentic applications**, precisely where memory forecasts currently assume explosive growth.\n\n- If\n**That is large enough to change the memory cycle given expanding effective memory.**", "url": "https://wpnews.pro/news/ai-memory-prices-and-the-engineering-response", "canonical_source": "https://paulkedrosky.com/ai-memory-prices-and-the-engineering-response/", "published_at": "2026-09-02 14:26:11+00:00", "updated_at": "2026-09-03 00:23:21.159662+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-infrastructure", "ai-research"], "entities": ["Google", "Semicon Taiwan"], "alternates": {"html": "https://wpnews.pro/news/ai-memory-prices-and-the-engineering-response", "markdown": "https://wpnews.pro/news/ai-memory-prices-and-the-engineering-response.md", "text": "https://wpnews.pro/news/ai-memory-prices-and-the-engineering-response.txt", "jsonld": "https://wpnews.pro/news/ai-memory-prices-and-the-engineering-response.jsonld"}}