The Current State of AI Chips
NVIDIA and Google are racing to dominate the AI chip market, with NVIDIA's Blackwell and Rubin architectures and Google's TPU v8 and Ironwood v7 leading the shift toward energy-efficient 'tokens per w…
NVIDIA and Google are racing to dominate the AI chip market, with NVIDIA's Blackwell and Rubin architectures and Google's TPU v8 and Ironwood v7 leading the shift toward energy-efficient 'tokens per w…
Nvidia Corp. has notified customers that AI server prices will rise more than 15% for systems shipping early next year, driven by memory costs that now account for about 25% of a high-end rack's build…
Marin 535B-A23B, a 535B-parameter model with 23B active parameters, began training this week in a fully open process, according to Percy Liang's announcement on X. The run will use 18.75T tokens on 11…
NVIDIA has identified OpenAI as one of its Lighthouse model builders using the GB200 NVL72 rack-scale system at data-center scale for next-generation model training and production inference. The confi…
Red Hat published a post arguing that 'the CPU is back' for LLM inference, citing an Intel and Georgia Tech paper that found CPU-side tool processing accounts for 50–90% of total latency in agentic wo…
NVIDIA Dynamo, an Apache-2.0 datacenter-scale inference orchestration layer from NVIDIA, sits above vLLM, SGLang, and TensorRT-LLM to coordinate multi-GPU clusters, with published benchmarks showing u…
Nvidia's Vera Rubin NVL72 delivers 5.4x performance per megawatt and 5x performance per dollar over GB200 NVL72 on DeepSeek R1 inference, according to early engineering samples from CoreWeave. The sec…
Nvidia Corp. revealed performance benchmarks for its next-generation Vera Rubin platform, showing a 10x improvement in tokens per watt on DeepSeek's R1 model compared to the previous GB200 NVL72 syste…
Nvidia partnered with Mizuho Financial Group on July 15, 2026, to build AI factories using Nvidia's confidential computing technology for secure private AI environments in banking, focusing on fraud d…
ASUS gave a tour of its thermal testing lab in Taiwan, where it tests AI servers in environmental chambers that simulate years of thermal stress in compressed timeframes. The lab includes a walk-in ch…
NVIDIA announced that its full-stack inference and training optimizations, including the GB200 NVL72 rack-scale system and NVIDIA Dynamo software, can maximize AI factory energy efficiency, reducing t…
Tesla filed a trademark for 'Megapod,' a modular AI data center hardware system, signaling plans to enter a market dominated by Nvidia's GB200 NVL72 and DGX SuperPOD. The product would bundle servers,…
NVIDIA's Blackwell platform swept MLPerf Training 6.0, achieving the fastest training times across all seven benchmarks, scaling to 8,192 GPUs on DeepSeek-V3 671B, and delivering up to 1.6x performanc…
NVIDIA and SchedMD have released a new topology/block plugin for Slurm 23.11 that enables topology-aware job scheduling on NVIDIA GB200 NVL72 systems, allowing workloads to be aligned with NVLink doma…