Hot reload vLLM and sglang configs
A new open-source tool called trimtab lets operators change SGLang and vLLM scheduler settings live, cutting configuration changes from a 1-7 minute redeploy to about 15 ms, with zero dropped requests…
A new open-source tool called trimtab lets operators change SGLang and vLLM scheduler settings live, cutting configuration changes from a 1-7 minute redeploy to about 15 ms, with zero dropped requests…
A new technical blog post by Iaroslav Elistratov presents a visual guide to building a B200 attention kernel from scratch in CUDA and PTX, achieving 94.4% of FlashAttention-4 performance on 4K, 8K, an…
Wafer AI CEO Emilio Andere claims that AMD GPUs can match Nvidia's performance with proper software optimization, citing a July 2026 test where Wafer's optimization achieved approximately 80% of Nvidi…
VLLM's Decode Context Parallelism (DCP) enables higher concurrency and throughput for long-context agentic workloads by splitting the KV cache across GPUs, according to a new blog post from the vLLM t…
Nvidia's $500 billion financing platform, which lets AI startups and data center operators purchase H100s, B200s, and future Blackwell architectures through structured financing, raises serious concer…
Compute grew 106x since NVIDIA's P100 in 2016, while memory bandwidth grew only 10.9x, a 10x decline in bytes-per-FLOP that has left AI accelerators bandwidth-starved, according to a chip inventory of…
Nvidia has notified customers and supply chain partners of price increases exceeding 15% on AI-related GPU products, driven by rising high-bandwidth memory costs and surging AI demand. Server GPUs lik…
Z.ai and NVIDIA engineers trained GLM-5.2, a 744B-parameter mixture-of-experts model quantized to 4-bit NVFP4, with reinforcement learning to play Super Mario Bros., overcoming arithmetic, distributed…
NVIDIA's B200 GPU has 180 GB of HBM3e memory with up to 8 TB/s bandwidth, while Micron's 245.76 TB 6600 ION SSD, shipping in May 2026, holds over a thousand times more data but with far lower bandwidt…
The Commodity Futures Trading Commission (CFTC) is seeking public input on futures contracts for AI computing power, with a draft request for comment sent to the White House's Office of Management and…
The Commodity Futures Trading Commission has filed a Request for Comment on derivatives contracts tied to computing capacity, seeking feedback on cash market liquidity, manipulation risks, customer pr…
Researchers introduced PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization, measuring functional correctness, ta…
Hyperscalers may face massive premiums for natural gas power as AI infrastructure demands unprecedented energy density, potentially driving up the cost per token for AI inference and forcing a pivot t…
Open-weight AI models have closed the performance gap with closed-source APIs, with specialized 7B to 30B models now outperforming older 175B models due to improved data quality and Mixture-of-Experts…
Nvidia is pursuing a $500 billion market valuation target by integrating its InfiniBand networking, CUDA software, and Blackwell architecture into a proprietary full-stack ecosystem, creating high swi…
An audit by computer science researchers finds that NVIDIA's Blackwell Ultra GPU (B300) nominally supports INT8 tensor-core compute at a 30:1 FP8-to-INT8 ratio, but the format is undeployable by defau…
CME Group, the Chicago-based exchange, announced it will launch futures contracts tied to the cost of renting Nvidia's H100 and B200 AI chips starting October 5, pending regulatory approval. The contr…
NVIDIA's flagship workstation GPU, the RTX PRO 6000 "Blackwell," now retails for $16,000, making it the most expensive non-server GPU the company sells. The price hike is driven by memory shortages, a…
CME Group plans to launch Silicon Data H100 Rental Index Futures and Silicon Data B200 Rental Index Futures on October 5, pending regulatory review, creating the first standardized financial instrumen…
CME Group will launch two compute futures contracts on October 5, 2026, enabling AI companies and investors to hedge against volatile GPU rental costs. The contracts, developed with Silicon Data, trac…