{"slug": "nemotron-3-diarization", "title": "Nemotron 3 Diarization", "summary": "NVIDIA released Nemotron 3 Diarization, a 100M-parameter open speaker-diarization model that ranks #1 on VoiceArena's Diarization-Bench leaderboard with a 14.72% Diarization Error Rate. The model supports up to eight speakers in real time, cutting DER by an average of 41% relative against NVIDIA's previous four-speaker streaming baseline at 1.04-second latency, and ships under the openmdw-1.1 license for commercial use with GGUF, ONNX, CoreML, and LiteRT community ports. NVIDIA reported 26,000 downloads in five days.", "body_md": "# Nemotron 3 Diarization\n\nedge\n**Who spoke when, in 107 megabytes.** Nemotron 3 Diarization is NVIDIA’s open speaker-diarization model: a 100M-parameter Sortformer descendant that decides which of up to eight speakers is talking, in real time, and ranks #1 on VoiceArena’s Diarization-Bench leaderboard at a 14.72% Diarization Error Rate. Against NVIDIA’s previous four-speaker streaming baseline it cuts DER by an average of 41% relative at 1.04-second latency, with the advantage widening as speaker counts rise - the eight-speaker support is what makes it a meetings model rather than a phone-call model.\n\n**Streaming and offline in one checkpoint.** One checkpoint covers both modes: input buffer as low as 80ms for latency-critical uses (the recommended floor is 0.32s), 30.4s buffer for offline-quality transcripts, output frames in 10ms multiples, and chunked inference with no duration limit. It produces speaker activity and timestamps, not words: chain it with a timestamped ASR model (Nemotron ASR 3.5 or Parakeet TDT 0.6B v3) and word-to-speaker mapping gives you attributed transcripts, the piece every local meeting-assistant pipeline was missing.\n\n**What runs where.** Two routes. The NeMo path wants Linux with an NVIDIA Ampere-or-newer GPU. The interesting one for this audience is NeMo-Speech.cpp, NVIDIA’s C++ runtime (`nemo-speech diarize meeting.wav`, or `--diarize` to tag a transcript with word-level speakers), plus community ports: a GGUF at 198.7 MB bf16 / 106.7 MB q8_0, ONNX, CoreML, and LiteRT builds, which put real-time diarization on CPUs and phones, not just CUDA boxes. License is openmdw-1.1, explicitly ready for commercial use, which is the difference between a demo and a pipeline component. 26,000 downloads in five days say the meeting-tools crowd noticed.\n\n- 100M\n- openmdw 1.1\n- 🇺🇸 USA\n- Sep 2026\n\n## Related models\n\nSave your hardware and every model page answers the real question: will it run on *your* machine, and how fast?\n\n[Join free - save your rig →](https://tokenstead.ai/login?return_to=%2Fonboarding)\n\n## Run it locally\n\nPer-quant memory needs and a static \"can you run it?\" reference - no rig entry required\n\n### The reference hardware\n\n22 reference configs, drawn in-house. Scroll for more.\n\n### Can you run it? - reference rigs\n\n| Rig | Q8_0 | BF16 | \n|---|---|---|\n| 4x H100 80GB (320GB) | fast 67000.0t/s | fast 36850.0t/s | \n| NVIDIA DGX Station 748GB | fast 40000.0t/s | fast 22000.0t/s | \n| 8x RTX 3090 rack (192GB) | fast 37448.0t/s | fast 20596.4t/s | \n| 4x RTX 5090 (128GB) | fast 35840.0t/s | fast 19712.0t/s | \n| AMD Instinct MI300X (192GB) | fast 26624.0t/s | fast 14643.2t/s | \n| 4x RTX 4090 (96GB) | fast 20160.0t/s | fast 11088.0t/s | \n| 2x RTX 5090 (64GB) | fast 17920.0t/s | fast 9856.0t/s | \n| 2x RTX 3090 (48GB) | fast 9362.0t/s | fast 5149.1t/s | \n| Single RTX 5090 (32GB) | fast 8960.0t/s | fast 4928.0t/s | \n| RTX PRO 6000 Blackwell (96GB) | fast 8960.0t/s | fast 4928.0t/s | \n| Mac Studio M4 Ultra 192GB | fast 5956.4t/s | fast 3276.0t/s | \n| Mac Studio M4 Ultra 512GB | fast 5956.4t/s | fast 3276.0t/s | \n| Single RTX 4090 (24GB) | fast 5040.0t/s | fast 2772.0t/s | \n| MacBook Pro M5 Max 128GB | fast 3349.1t/s | fast 1842.0t/s | \n| Single GTX 1080 Ti (11GB) | fast 2420.0t/s | fast 1331.0t/s | \n| Dual EPYC 9004 + 768GB DDR5-4800 | fast 2304.0t/s | fast 1267.2t/s | \n| DGX Spark 128GB unified | fast 1365.0t/s | fast 750.8t/s | \n| Ryzen AI Max+ 395 128GB | fast 1280.0t/s | fast 704.0t/s | \n| Jetson AGX Orin 64GB | fast 1024.0t/s | fast 563.2t/s | \n| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | fast 1024.0t/s | fast 563.2t/s | \n| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | fast 768.0t/s | fast 422.4t/s | \n| NVIDIA Jetson Orin NX 16GB | fast 512.0t/s | fast 281.6t/s | \n\nFit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.\n\n## Download options\n\n## Or run it in the cloud\n\n        No per-token API provider pricing tracked for Nemotron 3 Diarization yet.\n        For flagship list prices, see the\n        [calculator](https://tokenstead.ai/calculator).\n      \n\n## Inference cost over time\n\nData accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.", "url": "https://wpnews.pro/news/nemotron-3-diarization", "canonical_source": "https://tokenstead.ai/models/nemotron-3-diarization", "published_at": "2026-09-28 12:42:23+00:00", "updated_at": "2026-09-28 12:49:30.995979+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-products", "ai-tools", "natural-language-processing"], "entities": ["NVIDIA", "Nemotron 3 Diarization", "VoiceArena", "Diarization-Bench", "Nemotron ASR 3.5", "Parakeet TDT 0.6B v3", "NeMo-Speech.cpp", "NeMo"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/nemotron-3-diarization", "markdown": "https://wpnews.pro/news/nemotron-3-diarization.md", "text": "https://wpnews.pro/news/nemotron-3-diarization.txt", "jsonld": "https://wpnews.pro/news/nemotron-3-diarization.jsonld"}}