MaxKernel: Agentic Kernel Generation for TPUs
MaxKernel, a new agentic system presented by Google, uses large language models with real-time compiler feedback to generate high-performance custom kernels for Tensor Processing Units (TPUs), address…
MaxKernel, a new agentic system presented by Google, uses large language models with real-time compiler feedback to generate high-performance custom kernels for Tensor Processing Units (TPUs), address…
Maneshwar, the developer behind LiveReview, an AI code review tool, explains the differences between CPU, GPU, TPU, NPU, DPU, and QPU, arguing that 'computation' is not a single thing. He details how …
Google Cloud has integrated native TPU support into vLLM, the open-source LLM serving engine, to enable enterprise-grade precision for long-context multimodal embedding inference, targeting the Qwen3 …
Marvell Technology has granted Alphabet Inc.'s Google a warrant to buy up to 58.97 million Marvell shares at $206.58 each, worth roughly $12.18 billion, with vesting tied to Google's purchases of cust…
A developer documented their migration of a Cloud TPU workload from the deprecated Cloud TPU API to Compute Engine, moving a v6e-1 chip serving gemma-4-E2B-it under vLLM. The migration required mappin…
Engineers are integrating large language models into real-time data pipelines, facing challenges such as state synchronization and KV cache transfer bottlenecks. Optimizations using TensorRT-LLM and a…
Bernstein projects the global server market will exceed $1 trillion by 2028, driven by accelerating AI infrastructure spending, according to its Q2 2026 AI tracker report released on August 3, 2026. T…
Microsoft, Google, Amazon, and Meta are accelerating capital expenditures on AI infrastructure, with Microsoft expanding Azure data center leases and GPU fleets, Google committing to TPU buildouts, Am…
Amazon, Microsoft, and Alphabet have collectively guided their 2026 capital expenditure to roughly $725 billion, a 77% jump from the $410 billion projected for 2025, signaling massive demand for AI ch…
Nexus Data Centers is in advanced talks to secure $15 billion in bank financing for a Texas data center campus tied to Anthropic, with Alphabet's Google providing financial guarantees and supplying TP…
Google Chief Scientist Jeff Dean said at Y Combinator's Startup School 2026 that AI models have reached the capability of junior engineers, a prediction he made in May 2025, and predicted that by 2027…
Google released a microbenchmark suite on GitHub to help developers evaluate TPU performance by measuring network, compute, HBM, and host transfer capabilities. The suite establishes a Speed-of-Light …
Alphabet spent a record $44.9 billion in capital expenditures in Q2 2026, doubling from $22.5 billion a year earlier, driven by AI infrastructure investments, and reported its first-ever negative free…
Google's open-source TPU compiler, shipped as the Mosaic TPU dialect inside JAX, documents eight explicit memory spaces that programmers must manage, including vector memory (VMEM), scalar memory (SME…
Token dropping in mixture-of-experts (MoE) layers occurs when a capacity factor limits each expert's buffer, causing excess tokens to bypass the expert MLP and degrade model quality under production l…
XLA, the compiler behind JAX, TensorFlow, and PyTorch/XLA, optimizes array programs by freezing shapes, statically allocating buffers, and fusing operations against a global cost model, which makes it…
Google is reportedly developing 'Frozen v2,' a dedicated inference chip built specifically to run Gemini models, claiming a 6x to 10x increase in energy efficiency by hardcoding core Gemini operationa…
Google has started selling its custom Tensor Processing Units (TPUs) to customers for the first time, delivering them to customer data centers in the second quarter, CFO Anat Ashkenazi disclosed durin…
Google has released AlphaEvolve, an AI coding agent developed by Google DeepMind, on the Gemini Enterprise Agent Platform. The tool uses large language models to automatically generate and improve alg…
A surge in demand for non-Nvidia AI chips is reshaping hardware strategies as engineers optimize large language model deployments for alternative architectures due to H100 shortages and high costs. Th…