Apple Going All-In on AI Chips
Apple is accelerating development of its M7 Ultra chip, which will dramatically upgrade AI performance and bring it closer to dedicated AI accelerators like Nvidia Corp.'s Blackwell, according to Bloo…
Apple is accelerating development of its M7 Ultra chip, which will dramatically upgrade AI performance and bring it closer to dedicated AI accelerators like Nvidia Corp.'s Blackwell, according to Bloo…
Nvidia on July 1 announced software optimizations that boost token throughput by up to 5x for DeepSeek V4 on Blackwell systems, with throughput improvements reaching 20x compared to baseline configura…
Apple is developing the M7 Ultra chip with a potential 1.5TB unified memory capacity, targeting a 2028 release and aiming to deliver AI performance comparable to Nvidia's Blackwell accelerators, accor…
Apple's rumored M7 Ultra chip is targeting 1.5TB of unified memory and AI performance comparable to Nvidia's Blackwell architecture, according to leaks. The chip would represent a significant leap in …
A hyper-optimized, zero-dependency C/CUDA inference engine for the Qwen 3.6 35B model on RTX 5090 Blackwell GPUs achieves 13.4k tokens/sec prefill throughput at 2,048 context depth and 270+ tokens/sec…
NVIDIA's RTX 5090 offers 32GB of VRAM and 78% more memory bandwidth than the RTX 4090, but the extra performance is most noticeable for specific local LLM workloads such as running 32B models at Q4 wi…
NVIDIA, AMD, and Intel compete in the 2026 AI GPU market, with NVIDIA's Blackwell RTX 50-series, AMD's Radeon AI Pro R9700, and Intel's Arc Pro B70 targeting local LLM inference. VRAM capacity and mem…
Nvidia announced its RTX Spark ARM-based chip at Computex, positioning it as a new category for thin laptops and mini PCs rather than a replacement for x86 gaming desktops. The chip, made with MediaTe…
Parasail will deploy d-Matrix Corsair inference accelerators alongside NVIDIA Hopper and Blackwell GPUs in a heterogeneous inference system, targeting faster token generation for cloud customers. The …
NVIDIA researchers introduced Nonuniform Tensor Parallelism (NTP), a framework that dynamically adjusts tensor parallelism to maintain high Goodput during large-scale LLM training despite GPU interrup…
Wafer served GLM5.2 on AMD MI355X GPUs at 2626 tokens per second per node with over 2x lower cost than NVIDIA Blackwell, achieving 213 tok/s single stream. The company used MXFP4 quantization via AMD …
NVIDIA announced that its Confidential Computing technology for Blackwell GPUs achieves up to 98% of the inference performance of non-secure solutions, enabling hardware-rooted AI security without sig…
RadixArk released Miles, an open-source PyTorch-native framework for large-scale LLM reinforcement learning post-training. The framework integrates SGLang for rollout, NVIDIA Megatron-LM for training,…
NVIDIA's full-stack inference software, codesigned with its hardware, has reduced token costs by up to 5x on the DeepSeek V4 model in one month on the Blackwell platform. Companies like Baseten, Cogni…
CoreWeave has partnered with Swedish data center operator Conapto to add two Stockholm campuses to its European AI infrastructure, pushing its global data center total to 49 and active power capacity …
Nvidia is hiring for over 12 robotics and AI positions in China, focusing on its Omniverse platform and humanoid robotics, despite US export tensions. The company partnered with Unitree Robotics in Ju…
Multiverse Computing launched Pulsar 16B, an open-source reasoning model built on Nvidia's Nemotron architecture that achieves performance comparable to 30B-parameter models while using half the compu…
NVOC, an open-source Linux tool for overclocking NVIDIA GPUs, has reached version 0.3.0 with multi-GPU support, machine-readable JSON output, and improved memory overclocking for AI workloads. The too…
NVIDIA released the Nemotron 3 Ultra NVFP4 checkpoint, a quantized model that achieves up to 5.9x higher inference throughput than GLM-5.1 754B FP4 on decode-heavy workloads while matching BF16 accura…
A new book, 'Modern GPU Programming For MLSys', teaches GPU kernel optimization for machine learning systems, focusing on Blackwell architecture and techniques like GEMM and FlashAttention. Developed …