Fast Polynomial Transcendentals for LLMs
A new arXiv paper (2610.00049v1) reports that replacing native sigmoid, tanh, and SiLU activations with degree-3 or degree-4 bfloat16 polynomial programs improved complete training-step throughput on …
A new arXiv paper (2610.00049v1) reports that replacing native sigmoid, tanh, and SiLU activations with degree-3 or degree-4 bfloat16 polynomial programs improved complete training-step throughput on …
A September 29, 2026 arXiv paper reports that an LLM agent translating Triton kernels directly into NVIDIA PTX, a process the authors call "AI lowering," achieved 0.83x to 3.34x the performance of aut…
Nearly 90% of IT leaders have already experienced security incidents tied to AI pilot programs, according to industry data cited in a CIO guidance piece from HPE and NVIDIA. HPE and NVIDIA outlined th…
Eli Lilly chief information and digital officer Diogo Rau was named a TIME Executive of the Year after his team built LillyPod, an Nvidia-powered supercomputer with more than 1,000 Blackwell GPUs, in …
Nvidia's GeForce RTX 60 series, built on the architecture codenamed Rubin, is now targeting a 2028 launch window, according to leakers Kopite7kimi and Kepler_L2, a slip from earlier projections of 202…
PyTorch 2.14 shipped on September 15 with 2,995 commits from 487 contributors, adding fault-tolerant distributed training as a first-class framework concept via a rewritten NCCL backend (nccl2) with i…
SemiAnalysis reported that Nvidia's Vera Rubin NVL72 platform delivered up to 7x better token throughput per megawatt than Blackwell on its AgentX agentic inference benchmark, even on early pre-releas…
Pinterest expanded its partnership with Nvidia, pairing Nvidia's Blackwell GPUs and Dynamo inference framework with Pinterest's proprietary visual embeddings to achieve an 85x improvement in response …
Nvidia announced the RTX Pro 5500 Blackwell Workstation Edition, a rack-mountable professional GPU built on the Blackwell architecture with 84 GB of GDDR7 memory and a choice of air- or liquid-cooled …
Nvidia announced a partnership with unnamed Australian cloud and data center operators on Wednesday to build up to 2 gigawatts of AI computing capacity in Australia by 2027, with Nvidia supplying its …
NVIDIA confirmed it will bring DLSS 5 neural rendering to the GeForce RTX 40-Series "Ada Lovelace" graphics cards, after initially focusing on the RTX 50-Series "Blackwell." The company stated that DL…
Nvidia may acquire Hugging Face to create a unified AI stack from silicon to deployed models, according to an analysis. The potential merger would streamline model optimization, compute integration, a…
NVIDIA's summer intern work on chip-to-chip interconnect highlights that communication, not compute, is the bottleneck in scaling AI to trillion-parameter models, with bandwidth dropping from tens of …
NVIDIA confirmed at IFA that its upcoming N1X SoC for the RTX Spark PC platform will launch in two configurations: a top-end model with a 20-core Grace CPU and a Blackwell GPU with 6,144 CUDA cores, a…
Nvidia's Arm-based RTX Spark N1X chip will launch in October in two configurations for laptops and mini PCs, with the top-tier variant featuring a 6144-core Blackwell RTX GPU, 20-core Grace CPU, and 2…
A new technical blog post by Iaroslav Elistratov presents a visual guide to building a B200 attention kernel from scratch in CUDA and PTX, achieving 94.4% of FlashAttention-4 performance on 4K, 8K, an…
NVIDIA confirmed DLSS 5 will launch on September 3 with NBA 2K27 as its only supported game, and the neural renderer still costs 50-60% of native frame rate before Multi Frame Generation. The company …
Relace reports that search consumed more than half the tokens across 1,200 coding agent traces, and a dedicated retrieval model achieved 0.71 Recall@k while a compaction pass cut one real trace by 55 …
OpenAI and Broadcom co-developed Jalapeño, a custom inference ASIC that OpenAI claims beats NVIDIA's Blackwell on latency and performance per watt in early benchmarks. The chip, slated for deployment …
Abu Dhabi has operationalized multi-gigawatt AI compute hubs in 2026, with 5 GW of AI-only capacity and a PUE target of 1.10–1.18, shifting inference economics regionally and enabling data-local LLM s…