Evaluating Nemotron 3 Embed for agent memory
NVIDIA's Nemotron 3 Embed 1B embedding model took first place on every recall task in Zep's benchmark of 5,954 production queries, outperforming a 4B competitor with a third of the parameters and surp…
NVIDIA's Nemotron 3 Embed 1B embedding model took first place on every recall task in Zep's benchmark of 5,954 production queries, outperforming a 4B competitor with a third of the parameters and surp…
AMD has integrated an NVFP4 emulation pipeline into vLLM that enables AMD Instinct MI355 accelerators to serve standard NVFP4 quantized checkpoints directly, dequantizing weights to BF16 on-the-fly at…
ARCQuant, a new framework for Large Language Model inference, uses the NVFP4 numerical format to achieve up to 3x speedup on GPUs while maintaining accuracy comparable to full-precision baselines. The…
NVIDIA released the Nemotron 3 Ultra NVFP4 checkpoint, a quantized model that achieves up to 5.9x higher inference throughput than GLM-5.1 754B FP4 on decode-heavy workloads while matching BF16 accura…
NVIDIA's June 2026 DGX Spark update introduces automated four-node clustering via Cluster Assistant, enabling local inference of models up to 700B parameters. The update also delivers a 2.6x throughpu…
NVIDIA announced that its full-stack inference and training optimizations, including the GB200 NVL72 rack-scale system and NVIDIA Dynamo software, can maximize AI factory energy efficiency, reducing t…
A critical bug in Google's Gemma 4 causes it to malform tool calls under real load, affecting vLLM, llama.cpp, Ollama, and oobabooga. A developer open-sourced a diagnosis, repair, and experimental LoR…
NVIDIA released a guide showing how to optimize transformer-based models for low-precision training using Hopper and Blackwell GPUs, focusing on FP8 and NVFP4 formats. The method translates model conf…
NVIDIA released Nemotron 3 Ultra, a 550B-parameter hybrid Mamba-Transformer model with 55B active parameters, achieving up to 6x higher inference throughput than state-of-the-art LLMs while maintainin…
NVIDIA has introduced the NVFP4 training recipe in TransformerEngine, enabling 4-bit mixed-precision pre-training on Blackwell GPUs with no measurable accuracy loss compared to FP8 baselines. The reci…
NVIDIA Nemotron 3 Ultra, a 550-billion parameter open large language model, is now available for one-click deployment on Amazon SageMaker JumpStart. The model delivers 5x faster inference and up to 30…