Nemotron 3 Diarization NVIDIA released Nemotron 3 Diarization, a 100M-parameter open speaker-diarization model that ranks #1 on VoiceArena's Diarization-Bench leaderboard with a 14.72% Diarization Error Rate. The model supports up to eight speakers in real time, cutting DER by an average of 41% relative against NVIDIA's previous four-speaker streaming baseline at 1.04-second latency, and ships under the openmdw-1.1 license for commercial use with GGUF, ONNX, CoreML, and LiteRT community ports. NVIDIA reported 26,000 downloads in five days. Nemotron 3 Diarization edge Who spoke when, in 107 megabytes. Nemotron 3 Diarization is NVIDIA’s open speaker-diarization model: a 100M-parameter Sortformer descendant that decides which of up to eight speakers is talking, in real time, and ranks 1 on VoiceArena’s Diarization-Bench leaderboard at a 14.72% Diarization Error Rate. Against NVIDIA’s previous four-speaker streaming baseline it cuts DER by an average of 41% relative at 1.04-second latency, with the advantage widening as speaker counts rise - the eight-speaker support is what makes it a meetings model rather than a phone-call model. Streaming and offline in one checkpoint. One checkpoint covers both modes: input buffer as low as 80ms for latency-critical uses the recommended floor is 0.32s , 30.4s buffer for offline-quality transcripts, output frames in 10ms multiples, and chunked inference with no duration limit. It produces speaker activity and timestamps, not words: chain it with a timestamped ASR model Nemotron ASR 3.5 or Parakeet TDT 0.6B v3 and word-to-speaker mapping gives you attributed transcripts, the piece every local meeting-assistant pipeline was missing. What runs where. Two routes. The NeMo path wants Linux with an NVIDIA Ampere-or-newer GPU. The interesting one for this audience is NeMo-Speech.cpp, NVIDIA’s C++ runtime nemo-speech diarize meeting.wav , or --diarize to tag a transcript with word-level speakers , plus community ports: a GGUF at 198.7 MB bf16 / 106.7 MB q8 0, ONNX, CoreML, and LiteRT builds, which put real-time diarization on CPUs and phones, not just CUDA boxes. License is openmdw-1.1, explicitly ready for commercial use, which is the difference between a demo and a pipeline component. 26,000 downloads in five days say the meeting-tools crowd noticed. - 100M - openmdw 1.1 - 🇺🇸 USA - Sep 2026 Related models Save your hardware and every model page answers the real question: will it run on your machine, and how fast? Join free - save your rig → https://tokenstead.ai/login?return to=%2Fonboarding Run it locally Per-quant memory needs and a static "can you run it?" reference - no rig entry required The reference hardware 22 reference configs, drawn in-house. Scroll for more. Can you run it? - reference rigs | Rig | Q8 0 | BF16 | |---|---|---| | 4x H100 80GB 320GB | fast 67000.0t/s | fast 36850.0t/s | | NVIDIA DGX Station 748GB | fast 40000.0t/s | fast 22000.0t/s | | 8x RTX 3090 rack 192GB | fast 37448.0t/s | fast 20596.4t/s | | 4x RTX 5090 128GB | fast 35840.0t/s | fast 19712.0t/s | | AMD Instinct MI300X 192GB | fast 26624.0t/s | fast 14643.2t/s | | 4x RTX 4090 96GB | fast 20160.0t/s | fast 11088.0t/s | | 2x RTX 5090 64GB | fast 17920.0t/s | fast 9856.0t/s | | 2x RTX 3090 48GB | fast 9362.0t/s | fast 5149.1t/s | | Single RTX 5090 32GB | fast 8960.0t/s | fast 4928.0t/s | | RTX PRO 6000 Blackwell 96GB | fast 8960.0t/s | fast 4928.0t/s | | Mac Studio M4 Ultra 192GB | fast 5956.4t/s | fast 3276.0t/s | | Mac Studio M4 Ultra 512GB | fast 5956.4t/s | fast 3276.0t/s | | Single RTX 4090 24GB | fast 5040.0t/s | fast 2772.0t/s | | MacBook Pro M5 Max 128GB | fast 3349.1t/s | fast 1842.0t/s | | Single GTX 1080 Ti 11GB | fast 2420.0t/s | fast 1331.0t/s | | Dual EPYC 9004 + 768GB DDR5-4800 | fast 2304.0t/s | fast 1267.2t/s | | DGX Spark 128GB unified | fast 1365.0t/s | fast 750.8t/s | | Ryzen AI Max+ 395 128GB | fast 1280.0t/s | fast 704.0t/s | | Jetson AGX Orin 64GB | fast 1024.0t/s | fast 563.2t/s | | Epyc + 512GB DDR4-3200 + 2x RTX 3090 | fast 1024.0t/s | fast 563.2t/s | | Epyc + 512GB DDR4-2400 + 2x RTX 3090 | fast 768.0t/s | fast 422.4t/s | | NVIDIA Jetson Orin NX 16GB | fast 512.0t/s | fast 281.6t/s | Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast =20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark. Download options Or run it in the cloud No per-token API provider pricing tracked for Nemotron 3 Diarization yet. For flagship list prices, see the calculator https://tokenstead.ai/calculator . Inference cost over time Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.