The KV Cache Is the New Memory Wall
A Systematization of Knowledge paper, "The KV Cache Is the New Memory Wall" (arXiv:2609.30854) by Tejinder Singh of Dell Technologies, unifies five families of KV cache reduction techniques — quantiza…
A Systematization of Knowledge paper, "The KV Cache Is the New Memory Wall" (arXiv:2609.30854) by Tejinder Singh of Dell Technologies, unifies five families of KV cache reduction techniques — quantiza…
A scheduled inspection of a raised-floor plenum beneath a 2-megawatt AI cluster of 128 liquid-cooled racks, each holding eight NVIDIA H100 GPUs, uncovered zinc whiskers growing on galvanized floor sup…
Richael (dh8116), a Year 11 student in Auckland, built Soulor AI, an AI companion app with persistent memory and a five-perspective Simulation Mode, on a LoRA fine-tune of Qwen3-14B served with vLLM o…
IBM Research and CoreWeave Inc. co-designed identity management and workload isolation controls for agent workloads, extending IBM's internal identity systems into CoreWeave and using CoreWeave Sandbo…
Epoch AI reported that approximately 27.6 million H100-equivalent chips had been sold as of June 2026, a stock a September 2026 LessWrong assessment calculated could support between 24 million and 24 …
Federal prosecutors in Los Angeles arrested Greg Lui, 38, owner of Earthmade Computer Inc., on October 1, 2026, following a three-count indictment returned September 29 accusing him of moving more tha…
A new attributed record of AI compute infrastructure catalogs 83 facilities, 482 GPU clusters and 2,935 price observations, finding that official H100 rental prices range from $4.57 to $18.37 per acce…
Local Minutes fine-tuned Qwen3.5-4B with LoRA on 24,000 examples generated by Claude Opus 5.5, completing one training pass on a single rented H100 GPU in 1.6 hours for about $7, after Gemini 2.5 Flas…
Quail, an inference engine built by the QUery-Aware Inference Layer project, processes over 1 billion tokens per minute per H100 GPU on a multi-join AI-SQL query, more than 10x faster than the team's …
An engineering analysis of GPU infrastructure economics finds that owning an 8x H100 SXM5-class server carries roughly $310,000 in upfront capital plus about $37,500 in electricity and $4,300 per year…
Cast AI's 2026 State of Kubernetes Optimization Report found average GPU utilization across its fleet of 23,000 production clusters is just 5%, with AKS at 2%, EKS at 5%, and GKE at 6%, while the best…
A developer published a command-line guide for monitoring GPU utilization and temperature on remote servers, using nvidia-smi queries, cron-scheduled CSV logging, webhook-based temperature alerts, and…
LiquidAI released LFM2.5-VL-DSpark, a 279.5M-parameter speculative decoding drafter for its LFM2.5-VL-3B vision-language model that adds 8.9% to the target model's parameter count while delivering dec…
A new installment in the "How To Scale Your Model" series lays out a theory of sharded matrix multiplication for training large ML models across thousands of accelerators, using named-axis notation to…
Researchers introduced SWE-Serve, a benchmark of 53 repository-grounded tasks derived from recent production changes to SGLang, spanning six inference engineering families, to evaluate agents on produ…
An independent researcher pretrained a roughly 0.4B-parameter Bangla-first language model end-to-end in Rust for $164 in rented GPU time, documenting five defects in the Candle framework and three in …
Prism ML released Bonsai 2 27B, a ternary-quantized version of Qwen3.8-27B that stores weights as -1, 0, or +1 and fits in roughly 6 to 8.6GB while retaining 98.2% of the original FP16 model's benchma…
Anthropic open-sourced code that its Claude model wrote to optimize more than 30 open-source protein, genomics and sequence-modeling models, reporting an average speedup of about 4x with small precisi…
Ornn Data reported that the cheapest qualifying open-weight model on the Artificial Analysis Intelligence Index completes a task at roughly one fifth the cost of a comparable closed model, and that se…
Compute Assay launched ASSAY-1, a free, versioned specification, registry, and weekly report that defines the GPU-hour as a graded deliverable good, after finding a 3.8× price spread across H100 capac…