What Happened #
The most visible response to soaring memory prices—now perhaps more than 75% of a server bill of materials ( according to Google)—is higher semiconductor capex: new fabs, more HBM packaging, and an approaching flood of Chinese DRAM capacity.
As ever, too much analysis is static, naively extrapolating from the present and ignoring both engineering and supply responses. While the latter will come, the former is faster.
Google described at Semicon Taiwan this week how it is** redesigning its systems** to get more from scarce memory. Its new chips hold more data close at hand, share memory across thousands of processors, and compress the working memory used by AI models sixfold. Older, cheaper DRAM can also take over less demanding jobs. Combined, these changes create substantially more effective memory capacity without waiting for new fabs.
The upshot: These changes create a step change in more effective memory capacity, with no costly cleanroom required (even if that's coming). And Google isn't the only company taking these steps.
What It Means
The changes are multiplicative- If KV caches consume 15–40% of memory in a long-context inference workload.Compressing them sixfold reduces total memory requirements by roughly12–33%. - The same hardware could therefore support about 15–50% more inference work. - The largest gains appear in long-context and agentic applications, precisely where memory forecasts currently assume explosive growth.
- If That is large enough to change the memory cycle given expanding effective memory.