CCCL Runtime: A Modern C++ Runtime for CUDA
NVIDIA released CCCL Runtime, a modern C++ runtime for CUDA, as part of CUDA 13.2. The new APIs provide safer and more convenient abstractions for stream management, memory allocation, and kernel laun…
NVIDIA released CCCL Runtime, a modern C++ runtime for CUDA, as part of CUDA 13.2. The new APIs provide safer and more convenient abstractions for stream management, memory allocation, and kernel laun…
NVIDIA launched DAQIRI, a high-performance networking library for real-time AI data acquisition, enabling direct streaming from high-bandwidth detectors to GPU-accelerated processing. The technology a…
NVIDIA launched Halos for Robotics, a full-stack functional safety system for physical AI, extending its autonomous vehicle safety technology to industrial robots, humanoids, and autonomous mobile rob…
NVIDIA developer Norbert Juffa contributed a CUDA C++ implementation of the Haversine formula for fast great-circle distance calculations, leveraging the sinpi() and cospi() functions for improved per…
NVIDIA released XR AI, an open-source beta library for building intelligent agents on AR glasses and XR devices, integrating live camera and microphone streams with multimodal AI models and enterprise…
NVIDIA released a developer example for building a transaction foundation model using accelerated computing, demonstrating a near-50% lift in fraud detection Average Precision over an XGBoost baseline…
NVIDIA announced the ACE Game Agent SDK and new Unreal Engine 5 plugins at Unreal Fest 2026, enabling developers to build on-device AI companions for games. The SDK provides agentic, chat, and RAG API…
NVIDIA released a guide showing how to optimize transformer-based models for low-precision training using Hopper and Blackwell GPUs, focusing on FP8 and NVFP4 formats. The method translates model conf…
NVIDIA swept MLPerf Training v6.0 benchmarks, achieving the fastest training times at scale and highest per-accelerator performance across all tests, including new DeepSeek-V3 and GPT-OSS-20B workload…
NVIDIA released BioNeMo Recipes, a set of training recipes that use Low-Rank Adaptation (LoRA) to fine-tune large biological foundation models like ESM2-3B and Evo2-1B on a single workstation GPU. The…
NVIDIA introduced advanced fused MLP kernels for mixture-of-experts (MoE) models, built with the CuTe DSL, delivering 1.3x–2x kernel-level speedups and enabling sync-free MoE execution. The optimizati…
NVIDIA achieved leading agentic coding performance on the first agentic AI benchmark, AA-AgentPerf, delivering up to 20x better performance than previous generations. The benchmark, created by Artific…
NVIDIA and MiniMax released the M3 multimodal AI model on NVIDIA accelerated infrastructure, including Blackwell GPUs, enabling long-context reasoning and agentic workflows. The 428-billion parameter …
NVIDIA introduced intent-based security profiles in its Unified Fabric Manager for Quantum InfiniBand, enabling network administrators to configure multi-tenant fabric security with a single click. Th…
Google DeepMind's DiffusionGemma, optimized for NVIDIA platforms, generates text tokens in parallel rather than sequentially, achieving up to 1,000 tokens per second on a single NVIDIA H100 GPU. The m…
Battery energy storage systems (BESS) are becoming essential infrastructure for AI factories, which require power-dense, fast-changing electrical loads that differ from traditional data centers. Prope…
NVIDIA launched Enterprise Manageability for its DGX Spark and GB10 systems, providing IT teams with a complete operational framework covering the full lifecycle from provisioning to end-of-life retir…
NVIDIA released a workflow for converting FP8-quantized CLIP model checkpoints into TensorRT engines, enabling faster inference and higher GPU throughput for production deployment. The process involve…
NVIDIA introduced an automated, AI-driven research loop called Auto-FL within its NVIDIA FLARE platform to help researchers test and optimize federated learning strategies more efficiently. The system…
NVIDIA has introduced a new workflow using agent skills and Nemotron Speech to accelerate the evaluation of clinical automatic speech recognition (ASR) models, addressing the challenge of accurately r…