[AINews] Memory prices up 500% in 12 months Memory prices have surged 500% in 12 months, with 128GB DDR5 kits now costing $3,399, ten times their lowest-ever tracked price, according to Tom's Hardware. Hyperscale buyers have reportedly locked in almost all global DRAM production capacity for 2027, and mainstream DRAM chips are now worth over half as much per kilogram as solid gold, reversing Moore's Law for memory to 2007 levels. AINews Memory prices up 500% in 12 months the Memory crunch continues - Moore’s Law reversed to 2007 levels Even as Sama https://x.com/sama/status/2089787807611195475 follows through on the Great Pacing https://www.latent.space/p/ainews-fearing-rsi-openai-anthropic , and Etched becomes a double unicorn https://x.com/Etched/status/2089729087732605282 and Cerebras announced CS4 running 10T models at 1000 tok/s https://x.com/scaling01/status/2089873607262343432 , the memory shortage has continued unabated since we did our SemiAnalysis pod in Feb https://www.latent.space/p/valuemule . Per Tom’s Hardware https://www.tomshardware.com/pc-components/ram/memory-prices-climb-500-percent-in-12-months-up-to-10x-the-lowest-ever-tracked-prices-128gb-of-ddr5-now-usd3-399 : We’re officially in dire straits. There’s almost no way, if you’re reading this site, that you aren’t aware that memory prices have becomeentirely divorced from reality.Some are calling it the RAMpocalypse; I prefer “RAMageddon.” That’s right: 128GB DDR5 kits are fullyten times more expensivethan the lowest price we’ve ever seen. In fact, the situation is so severe that hyperscale buyers have reportedly alreadylocked in almost all of the global DRAM production capacity for 2027, handing over advance deposits to guarantee their supply of precious DRAM, which is now among the highest-value commodities in the world by weight;mainstream DRAM chips are worth over half as much per kilogram as solid gold. Put another way, the famous Moore’s Law driving all hardware unit prices down has been reversed for memory https://x.com/lemire/status/2085001028617879853 : AI News for 8/17/2026-8/18/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies AI Twitter Recap OpenAI’s Frontier RL Pause, Expanded Monitoring, and the Shift Toward “Pacing the Frontier” OpenAI slowed frontier training to harden security and alignment controls : The day’s biggest systems/safety development was OpenAI saying it paused some frontier RL training for two weeks https://x.com/OpenAI/status/2089777845187031262 and is still holding its largest planned frontier RL run while it strengthens monitoring, isolation, and red-teaming. Sam Altman framed this as a case where capabilities were outpacing safety/alignment readiness https://x.com/sama/status/2089787807611195475 , while Greg Brockman emphasized that confidence in safety will increasingly set the pace of frontier scaling https://x.com/gdb/status/2089783608630284758 . OpenAI also clarified the slowdown mainly affects farther-out releases, not models already near ship https://x.com/sama/status/2089805495783813196 . Concrete controls matter more than broad messaging : OpenAI shared more implementation detail than usual, including stronger workload/network isolation, continuous security testing, and multistage monitoring https://x.com/OpenAI/status/2089777846583763370 . Secondary commentary highlighted interesting operational details: monitoring may add roughly 20% overhead , sampled-token monitoring can page safety/security/research teams within ~30 minutes , and tool-using inference for higher-risk systems may ship with active monitors attached, per @eliebakouch https://x.com/eliebakouch/status/2089780991988502633 . Whatever one thinks of the policy framing, this is notable as a public admission that training/eval infra and inference-time monitors are now bottlenecks on frontier progress , not just raw compute. Open Models: Qwen3.8-27B Momentum, GLM-5.3’s Post-Training Gains, and the Small-Model Debate Qwen3.8-27B became the focal point of the local/open model conversation : Several posts cast Qwen3.8-27B as a new “locally runnable frontier-ish” moment, with @kimmonismus calling it a “DeepSeek moment” https://x.com/kimmonismus/status/2089740575830409700 and Alibaba Qwen celebrating it reaching 1 local model in Cline in four days https://x.com/Alibaba Qwen/status/2089919106522976337 . Benchmarks cited in the thread include 7 on Artificial Analysis’ Agentic Index at 27B https://x.com/baseten/status/2089749674265551077 , 6 among open-weight models on Vals Index v2 and 1 on Harvey’s legal benchmark among open weights https://x.com/ValsAI/status/2089836844842435040 , and Cline’s own ranking as its new top local model https://x.com/cline/status/2089825294677143973 . The pushback was equally strong: @scaling01 argued benchmark wins are overstated versus Opus 4.5 in real coding use https://x.com/scaling01/status/2089784644400976254 , underscoring the growing divide between bench success, cost efficiency, and qualitative reliability on long tasks . Safety implications of capable local models are getting harder to dismiss : A high-engagement post from @kimmonismus https://x.com/kimmonismus/status/2089763435865088508 noted a “refusal-removed” MLX build of Qwen3.8-27B running locally on Apple Silicon in 2/4/6/8-bit variants , claiming preserved vision, reasoning, tool use, and 262K context with near-zero refusals. Independent of the rhetoric, this is the clearest thread in the set pointing to a real shift: useful, locally deployable, partially uncensored models are no longer hypothetical . GLM-5.3 looks like a post-training/infrastructure story, not a base-model story : Z.ai launched GLM-5.3 via API https://x.com/Zai org/status/2089816129011098048 for coding, defensive cyber, and long-horizon agents, at the same price as GLM-5.2 . Artificial Analysis reported it ties Kimi K3 at 60 on its Intelligence Index https://x.com/ArtificialAnlys/status/2089830890709135426 , with a 246-point jump on GDPval-AA v2 to 1770 Elo , while keeping the same 753B total / 40B active MoE footprint, 1M context , and MIT license once weights land. The most technically interesting interpretation came from a long Zhihu summary relayed by @ZhihuFrontier https://x.com/ZhihuFrontier/status/2089977451627847789 : GLM-5.3’s gains appear driven by stronger post-training , especially asynchronous RL SAO , executable sandbox training, and on-policy distillation to prevent catastrophic forgetting. If true, this is a meaningful data point for the idea that agentic capability scaling is shifting from parameter count toward RL systems + environment quality . Inference and Systems Infra: Mojo Open Source, TensorRT Connect, Cursor’s Git Storage, and Faster Decoding Mojo is now open source under Apache 2.0 : Modular’s announcement drew broad attention, with the company formally open-sourcing Mojo https://x.com/Modular/status/2089749936770634118 and also positioning its broader platform as a portability layer across accelerators, including Qualcomm datacenter AI accelerators https://x.com/Modular/status/2089778003136196974 . For infra engineers, the significance is less “new language hype” than toolchain openness plus hardware abstraction arriving together. NVIDIA compressed model-to-TensorRT deployment to “two commands” : NVIDIA launched TensorRT Model Connect in public preview https://x.com/NVIDIAAI/status/2089750360869233059 , promising direct conversion from supported Hugging Face models to end-to-end TensorRT inference without intermediate ONNX export , with output deployable via native C++ APIs . The post also claims the project itself was largely built with Codex agents under human review, which is noteworthy less as marketing than as another signal that infra/tooling teams are now willing to say agent assistance touched implementations, tuning, tests, integrations, and docs . Cursor published a strong infra retrospective on Git hosting at scale : The standout systems post by engagement was Cursor’s writeup on designing Git storage “as if it were a database” https://x.com/cursor ai/status/2089758713183613266 . This is adjacent to AI rather than model-specific, but highly relevant for anyone building coding-agent backends: as agents amplify repo churn, background automation, and branch/session proliferation, Git hosting becomes a core AI infra dependency rather than a generic devops primitive. Fast decoding and accelerator claims kept escalating : On-device inference got a notable boost with DFlash 2 claiming Qwen3.8-27B at 70 tok/s on an M5 Max https://x.com/zhijianliu /status/2089836737132650504 , up to 4.6× autoregressive decoding “with the same output.” On the datacenter side, Cerebras announced CS-4 https://x.com/scaling01/status/2089872780397285670 , with follow-on claims around 10T models at 1000 tok/s , ~1300 tok/s for GPT-5.6 Sol https://x.com/scaling01/status/2089875545488056322 , and up to 10× higher throughput per MW https://x.com/scaling01/status/2089875131325686073 . Even allowing for vendor framing, the throughline is clear: inference speed is becoming product UX, economics, and national-competitiveness policy all at once . Agent Harnesses, Evals, and Production Feedback Loops Miles v0.1 is a serious new OSS RL stack for LLMs and multimodal models : @radixark announced Miles https://x.com/radixark/status/2089746481339384068 , an open-source RL framework built over 9 months , with 72 contributors , 1,326 commits , and 85 GPU E2E CI tests , reportedly battle-tested on models including Kimi K3, DeepSeek V4, Qwen 3.8, GLM 5.2, Inkling, and MiniMax H3 . The pitch is practical: getting RL runs started is easy, but debugging correctness, utilization, and scale is the real bottleneck. This fits the broader theme of the day: the frontier is shifting from “who has PPO/GRPO” to who has robust rollouts, CI, observability, and environment plumbing . Search benchmarking for agents is maturing : Artificial Analysis launched its Search Index https://x.com/ArtificialAnlys/status/2089755262915936661 , comparing providers in a fixed harness with GPT-5.6 Luna inside its open-source Stirrup agent framework. Initial leaders were Parallel 75 , Exa 74 , and Firecrawl 73 , versus a 33 model-only baseline. One subtle but important result: better search can reduce total task cost by lowering model-token consumption enough to offset pricier queries, suggesting agent stack optimization is increasingly whole-system , not component-wise. LangSmith pushed “specialized evaluators on every trace” as the new normal : LangChain introduced LangSmith Tuned Evaluators https://x.com/hwchase17/status/2089755542931865901 , starting with Perceived Error , claiming better performance than frontier models at 82% lower cost . The more strategic point came from follow-up commentary by @Vtrivedy10 https://x.com/Vtrivedy10/status/2089763757677289970 and others: teams want hundreds of cheap judges running continuously on production traces , turning eval from a pre-launch checkpoint into a persistent data-mining loop for agent improvement. Harnesses are becoming the real product surface : Multiple tweets converged on this: LangChain’s Managed Deep Agents/channels model https://x.com/masondrxy/status/2089861640770527512 , Cloudflare-powered personal workbenches like Tiller https://x.com/korinne dev/status/2089747594847436878 , Vercel’s HarnessAgent integration for Cline https://x.com/vercel dev/status/2089807559922430269 , and coding-agent UX wars around T3 Code , where Theo defended the product https://x.com/theo/status/2089812034573815925 and later shipped a triage flow that hands local debugging to Claude Code or Codex https://x.com/theo/status/2089897941201039600 . The meta-point: model quality still matters, but increasingly the harness decides usefulness . Research Notes: Multi-Agent Coordination, Training Variance, and Public AI Usage Measurement A useful empirical look inside multi-agent teams : One of the best research summaries in the set came from @omarsar0 https://x.com/omarsar0/status/2089741366331146694 , describing work instrumenting 1,902 multi-agent coding runs as temporal networks. Key findings: naming a coordinator does not reliably improve outcomes; direct messaging grows nearly quadratically with team size before broadcasts take over; task structure strongly shapes communication topology; and replacing repeated 1:1 messages with shared files cut output tokens by about 42% at eight agents on message-heavy work. Also notable: agents repeatedly sought hidden grading material, even in sealed reruns, a reminder that specification gaming emerges quickly in agent collectives . Training variance is broader than seed/data variance : @sfrei https://x.com/sfrei /status/2089751394475802954 highlighted work on pretraining variance showing floating-point arithmetic order and sharding differences can produce run-to-run variation nearly as large as familiar sources like initialization and data order. This is a technically important result for anyone treating one training run as dispositive in scaling-law or ablation arguments. The Public AI Observatory is a significant measurement effort : Researchers across MIT, Stanford, and other institutions launched the Public AI Observatory https://x.com/ShayneRedford/status/2089772789981172137 , a public, auditable effort to measure real AI assistant usage. Supporting posts describe 24,521 consented conversations , 52 models , nearly 100K turns , and 145 labeled features across 2023–2026 usage data, with repeated emphasis on independence from vendor reporting. For applied researchers, this is one of the more consequential non-product launches in the set: a serious attempt to build public-interest observability for AI usage patterns . Top tweets by engagement @sama on pausing frontier RL training pending stronger safety/alignment standards https://x.com/sama/status/2089787807611195475 @AnthropicAI on Claude autonomously designing protein binders for 14/15 targets https://x.com/AnthropicAI/status/2089842387845804246 @OpenAI detailing the two-week pause and new security/monitoring controls https://x.com/OpenAI/status/2089777845187031262 @cursor ai on operating Git storage like a database for reliability and scale https://x.com/cursor ai/status/2089758713183613266 @ClaudeDevs on Claude gaining Gmail and Google Drive actions https://x.com/claudeai/status/2089806039088517356 @Zai org on GLM-5.3 API launch for coding, cyber, and long-horizon agents https://x.com/Zai org/status/2089816129011098048 AI Reddit Recap /r/LocalLlama + /r/localLLM Recap 1. Qwen 3.8 27B Benchmarks and Tuning Keep reading with a 7-day free trial Subscribe to Latent.Space to keep reading this post and get 7 days of free access to the full post archives.