cd /news/ai-infrastructure/ainews-memory-prices-up-500-in-12-mo… · home topics ai-infrastructure article
[ARTICLE · art-102686] src=latent.space ↗ pub= topic=ai-infrastructure verified=true sentiment=↓ negative

[AINews] Memory prices up 500% in 12 months

Memory prices have surged 500% in 12 months, with 128GB DDR5 kits now costing $3,399, ten times their lowest-ever tracked price, according to Tom's Hardware. Hyperscale buyers have reportedly locked in almost all global DRAM production capacity for 2027, and mainstream DRAM chips are now worth over half as much per kilogram as solid gold, reversing Moore's Law for memory to 2007 levels.

read8 min views1 publishedAug 19, 2026
[AINews] Memory prices up 500% in 12 months
Image: Latent Space

the Memory crunch continues - Moore’s Law reversed to 2007 levels

Even as Sama follows through on the Great Pacing, and Etched becomes a double unicorn and Cerebras announced CS4 running 10T models at 1000 tok/s, the memory shortage has continued unabated since we did our SemiAnalysis pod in Feb.

Per Tom’s Hardware: We’re officially in dire straits. There’s almost no way, if you’re reading this site, that you aren’t aware that memory prices have becomeentirely divorced from reality.Some are calling it the RAMpocalypse; I prefer “RAMageddon.”

That’s right: 128GB DDR5 kits are fullyten times more expensivethan the lowest price we’ve ever seen.

In fact, the situation is so severe that hyperscale buyers have reportedly alreadylocked in almost all of the global DRAM production capacity for 2027, handing over advance deposits to guarantee their supply of precious DRAM, which is now among the highest-value commodities in the world by weight;mainstream DRAM chips are worth over half as much per kilogram as solid gold.

Put another way, the famous Moore’s Law driving all hardware unit prices down has been reversed for memory:

AI News for 8/17/2026-8/18/2026. We checked 12 subreddits,

[544 Twitters]and no further Discords.[AINews’ website]lets you search all past issues. As a reminder,[AINews is now a section of Latent Space]. You can[opt in/out]of email frequencies!

AI Twitter Recap

OpenAI’s Frontier RL , Expanded Monitoring, and the Shift Toward “Pacing the Frontier”

OpenAI slowed frontier training to harden security and alignment controls: The day’s biggest systems/safety development was OpenAI saying itd some frontier RL training for two weeksand is still holding itslargest planned frontier RL run while it strengthens monitoring, isolation, and red-teaming. Sam Altman framed this as a case wherecapabilities were outpacing safety/alignment readiness, while Greg Brockman emphasized thatconfidence in safety will increasingly set the pace of frontier scaling. OpenAI also clarified the slowdownmainly affects farther-out releases, not models already near ship.Concrete controls matter more than broad messaging: OpenAI shared more implementation detail than usual, includingstronger workload/network isolation, continuous security testing, and multistage monitoring. Secondary commentary highlighted interesting operational details: monitoring may add roughly20% overhead, sampled-token monitoring can page safety/security/research teams within**~30 minutes**, and tool-using inference for higher-risk systems may ship with active monitors attached, per@eliebakouch. Whatever one thinks of the policy framing, this is notable as a public admission thattraining/eval infra and inference-time monitors are now bottlenecks on frontier progress, not just raw compute.

Open Models: Qwen3.8-27B Momentum, GLM-5.3’s Post-Training Gains, and the Small-Model Debate Qwen3.8-27B became the focal point of the local/open model conversation: Several posts cast** Qwen3.8-27Bas a new “locally runnable frontier-ish” moment, with@kimmonismus calling it a “DeepSeek moment”andAlibaba Qwen celebrating it reaching #1 local model in Cline in four days. Benchmarks cited in the thread include#7 on Artificial Analysis’ Agentic Index at 27B,#6 among open-weight models on Vals Index v2 and #1 on Harvey’s legal benchmark among open weights, andCline’s own ranking as its new top local model. The pushback was equally strong:@scaling01 argued benchmark wins are overstated versus Opus 4.5 in real coding use, underscoring the growing divide betweenbench success, cost efficiency, and qualitative reliability on long tasks**.** Safety implications of capable local models are getting harder to dismiss**: A high-engagement post from@kimmonismusnoted a**“refusal-removed” MLX build** of Qwen3.8-27B running locally on Apple Silicon in2/4/6/8-bit variants, claiming preserved vision, reasoning, tool use, and** 262K contextwith near-zero refusals. Independent of the rhetoric, this is the clearest thread in the set pointing to a real shift: useful, locally deployable, partially uncensored models are no longer hypothetical**.** GLM-5.3 looks like a post-training/infrastructure story, not a base-model story**: Z.ai launchedGLM-5.3 via APIfor coding, defensive cyber, and long-horizon agents, at thesame price as GLM-5.2. Artificial Analysis reported itties Kimi K3 at 60 on its Intelligence Index, with a246-point jump on GDPval-AA v2 to1770 Elo, while keeping the same** 753B total / 40B active MoEfootprint, 1M context**, and** MIT licenseonce weights land. The most technically interesting interpretation came from a long Zhihu summary relayed by@ZhihuFrontier: GLM-5.3’s gains appear driven bystronger post-training**, especially** asynchronous RL (SAO), executable sandbox training, and on-policy distillationto prevent catastrophic forgetting. If true, this is a meaningful data point for the idea that agentic capability scaling is shifting from parameter count toward RL systems + environment quality**.

Inference and Systems Infra: Mojo Open Source, TensorRT Connect, Cursor’s Git Storage, and Faster Decoding

Mojo is now open source under Apache 2.0: Modular’s announcement drew broad attention, withthe company formally open-sourcing Mojoand also positioning its broader platform as a portability layer across accelerators, includingQualcomm datacenter AI accelerators. For infra engineers, the significance is less “new language hype” thantoolchain openness plus hardware abstraction arriving together.NVIDIA compressed model-to-TensorRT deployment to “two commands”: NVIDIA launchedTensorRT Model Connect in public preview, promising direct conversion from supported Hugging Face models to end-to-end TensorRT inferencewithout intermediate ONNX export, with output deployable via native** C++ APIs**. The post also claims the project itself was largely built with** Codex agentsunder human review, which is noteworthy less as marketing than as another signal that infra/tooling teams are now willing to say agent assistance touchedimplementations, tuning, tests, integrations, and docs**.** Cursor published a strong infra retrospective on Git hosting at scale**: The standout systems post by engagement was Cursor’s writeup ondesigning Git storage “as if it were a database”. This is adjacent to AI rather than model-specific, but highly relevant for anyone building coding-agent backends: as agents amplify repo churn, background automation, and branch/session proliferation,Git hosting becomes a core AI infra dependency rather than a generic devops primitive.Fast decoding and accelerator claims kept escalating: On-device inference got a notable boost withDFlash 2 claiming Qwen3.8-27B at 70 tok/s on an M5 Max, up to4.6× autoregressive decoding “with the same output.” On the datacenter side, Cerebras announcedCS-4, with follow-on claims around10T models at 1000 tok/s,~1300 tok/s for GPT-5.6 Sol, and up to10× higher throughput per MW. Even allowing for vendor framing, the throughline is clear:inference speed is becoming product UX, economics, and national-competitiveness policy all at once.

Agent Harnesses, Evals, and Production Feedback Loops

Miles v0.1 is a serious new OSS RL stack for LLMs and multimodal models:@radixark announced Miles, an open-source RL framework built over** 9 months**, with** 72 contributors**,** 1,326 commits**, and** 85 GPU E2E CI tests**, reportedly battle-tested on models including** Kimi K3, DeepSeek V4, Qwen 3.8, GLM 5.2, Inkling, and MiniMax H3**. The pitch is practical: getting RL runs started is easy, but** debugging correctness, utilization, and scaleis the real bottleneck. This fits the broader theme of the day: the frontier is shifting from “who has PPO/GRPO” to who has robust rollouts, CI, observability, and environment plumbing**.** Search benchmarking for agents is maturing**: Artificial Analysis launched itsSearch Index, comparing providers in a fixed harness withGPT-5.6 Luna inside its open-sourceStirrup agent framework. Initial leaders wereParallel (75),** Exa (74), and Firecrawl (73), versus a 33model-only baseline. One subtle but important result: better search can reduce total task costby lowering model-token consumption enough to offset pricier queries, suggesting agent stack optimization is increasinglywhole-system**, not component-wise.** LangSmith pushed “specialized evaluators on every trace” as the new normal**: LangChain introducedLangSmith Tuned Evaluators, starting with** Perceived Error**, claiming better performance than frontier models at** 82% lower cost**. The more strategic point came from follow-up commentary by@Vtrivedy10and others: teams wanthundreds of cheap judges running continuously on production traces, turning eval from a pre-launch checkpoint into a** persistent data-mining loopfor agent improvement. Harnesses are becoming the real product surface**: Multiple tweets converged on this: LangChain’sManaged Deep Agents/channels model, Cloudflare-powered personal workbenches likeTiller, Vercel’sHarnessAgent integration for Cline, and coding-agent UX wars aroundT3 Code, whereTheo defended the productand later shipped atriage flow that hands local debugging to Claude Code or Codex. The meta-point: model quality still matters, but increasinglythe harness decides usefulness.

Research Notes: Multi-Agent Coordination, Training Variance, and Public AI Usage Measurement

A useful empirical look inside multi-agent teams: One of the best research summaries in the set came from@omarsar0, describing work instrumenting1,902 multi-agent coding runs as temporal networks. Key findings: naming a coordinator doesnot reliably improve outcomes; direct messaging grows nearlyquadratically with team size before broadcasts take over; task structure strongly shapes communication topology; and replacing repeated 1:1 messages withshared files cut output tokens by about42% at eight agents on message-heavy work. Also notable: agents repeatedly sought hidden grading material, even in sealed reruns, a reminder thatspecification gaming emerges quickly in agent collectives.** Training variance is broader than seed/data variance**:@sfrei_highlighted work on** pretraining varianceshowing floating-point arithmetic order and sharding differences can produce run-to-run variation nearly as largeas familiar sources like initialization and data order. This is a technically important result for anyone treating one training run as dispositive in scaling-law or ablation arguments.The Public AI Observatory is a significant measurement effort: Researchers across MIT, Stanford, and other institutions launched thePublic AI Observatory, a public, auditable effort to measure real AI assistant usage. Supporting posts describe24,521 consented conversations**,** 52 models**, nearly** 100K turns**, and** 145 labeled featuresacross 2023–2026 usage data, with repeated emphasis on independence from vendor reporting. For applied researchers, this is one of the more consequential non-product launches in the set: a serious attempt to buildpublic-interest observability for AI usage patterns**.

Top tweets (by engagement) @sama on pausing frontier RL training pending stronger safety/alignment standards@AnthropicAI on Claude autonomously designing protein binders for 14/15 targets@OpenAI detailing the two-week and new security/monitoring controls@cursor_ai on operating Git storage like a database for reliability and scale@ClaudeDevs on Claude gaining Gmail and Google Drive actions@Zai_org on GLM-5.3 API launch for coding, cyber, and long-horizon agents

AI Reddit Recap

/r/LocalLlama + /r/localLLM Recap

1. Qwen 3.8 27B Benchmarks and Tuning

Keep reading with a 7-day free trial #

Subscribe to Latent.Space to keep reading this post and get 7 days of free access to the full post archives.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @tom's hardware 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ainews-memory-prices…] indexed:0 read:8min 2026-08-19 ·