cd /news/artificial-intelligence/ai-memory-prices-and-the-engineering… · home topics artificial-intelligence article
[ARTICLE · art-119632] src=paulkedrosky.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

AI, Memory Prices, and the Engineering Response

Google said at Semicon Taiwan that it is redesigning its systems to cope with memory prices that now account for more than 75% of a server bill of materials, according to Google. The company's new chips hold more data close, share memory across thousands of processors, and compress AI working memory sixfold, creating substantially more effective memory capacity without waiting for new fabs. Compressing KV caches, which consume 15–40% of memory in long-context inference, could reduce total memory requirements by 12–33% and support 15–50% more inference work.

read1 min views1 publishedSep 2, 2026
AI, Memory Prices, and the Engineering Response
Image: Paulkedrosky (auto-discovered)

What Happened #

The most visible response to soaring memory prices—now perhaps more than 75% of a server bill of materials ( according to Google)—is higher semiconductor capex: new fabs, more HBM packaging, and an approaching flood of Chinese DRAM capacity.

As ever, too much analysis is static, naively extrapolating from the present and ignoring both engineering and supply responses. While the latter will come, the former is faster.

Google described at Semicon Taiwan this week how it is** redesigning its systems** to get more from scarce memory. Its new chips hold more data close at hand, share memory across thousands of processors, and compress the working memory used by AI models sixfold. Older, cheaper DRAM can also take over less demanding jobs. Combined, these changes create substantially more effective memory capacity without waiting for new fabs.

The upshot: These changes create a step change in more effective memory capacity, with no costly cleanroom required (even if that's coming). And Google isn't the only company taking these steps.

What It Means

The changes are multiplicative- If KV caches consume 15–40% of memory in a long-context inference workload.Compressing them sixfold reduces total memory requirements by roughly12–33%. - The same hardware could therefore support about 15–50% more inference work. - The largest gains appear in long-context and agentic applications, precisely where memory forecasts currently assume explosive growth.

  • If That is large enough to change the memory cycle given expanding effective memory.
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-memory-prices-and…] indexed:0 read:1min 2026-09-02 ·