edge Who spoke when, in 107 megabytes. Nemotron 3 Diarization is NVIDIA’s open speaker-diarization model: a 100M-parameter Sortformer descendant that decides which of up to eight speakers is talking, in real time, and ranks #1 on VoiceArena’s Diarization-Bench leaderboard at a 14.72% Diarization Error Rate. Against NVIDIA’s previous four-speaker streaming baseline it cuts DER by an average of 41% relative at 1.04-second latency, with the advantage widening as speaker counts rise - the eight-speaker support is what makes it a meetings model rather than a phone-call model.
Streaming and offline in one checkpoint. One checkpoint covers both modes: input buffer as low as 80ms for latency-critical uses (the recommended floor is 0.32s), 30.4s buffer for offline-quality transcripts, output frames in 10ms multiples, and chunked inference with no duration limit. It produces speaker activity and timestamps, not words: chain it with a timestamped ASR model (Nemotron ASR 3.5 or Parakeet TDT 0.6B v3) and word-to-speaker mapping gives you attributed transcripts, the piece every local meeting-assistant pipeline was missing.
What runs where. Two routes. The NeMo path wants Linux with an NVIDIA Ampere-or-newer GPU. The interesting one for this audience is NeMo-Speech.cpp, NVIDIA’s C++ runtime (nemo-speech diarize meeting.wav, or --diarize to tag a transcript with word-level speakers), plus community ports: a GGUF at 198.7 MB bf16 / 106.7 MB q8_0, ONNX, CoreML, and LiteRT builds, which put real-time diarization on CPUs and phones, not just CUDA boxes. License is openmdw-1.1, explicitly ready for commercial use, which is the difference between a demo and a pipeline component. 26,000 downloads in five days say the meeting-tools crowd noticed.
- 100M
- openmdw 1.1
- 🇺🇸 USA
- Sep 2026
Related models #
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Run it locally #
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | Q8_0 | BF16 |
|---|---|---|
| 4x H100 80GB (320GB) | fast 67000.0t/s | fast 36850.0t/s |
| NVIDIA DGX Station 748GB | fast 40000.0t/s | fast 22000.0t/s |
| 8x RTX 3090 rack (192GB) | fast 37448.0t/s | fast 20596.4t/s |
| 4x RTX 5090 (128GB) | fast 35840.0t/s | fast 19712.0t/s |
| AMD Instinct MI300X (192GB) | fast 26624.0t/s | fast 14643.2t/s |
| 4x RTX 4090 (96GB) | fast 20160.0t/s | fast 11088.0t/s |
| 2x RTX 5090 (64GB) | fast 17920.0t/s | fast 9856.0t/s |
| 2x RTX 3090 (48GB) | fast 9362.0t/s | fast 5149.1t/s |
| Single RTX 5090 (32GB) | fast 8960.0t/s | fast 4928.0t/s |
| RTX PRO 6000 Blackwell (96GB) | fast 8960.0t/s | fast 4928.0t/s |
| Mac Studio M4 Ultra 192GB | fast 5956.4t/s | fast 3276.0t/s |
| Mac Studio M4 Ultra 512GB | fast 5956.4t/s | fast 3276.0t/s |
| Single RTX 4090 (24GB) | fast 5040.0t/s | fast 2772.0t/s |
| MacBook Pro M5 Max 128GB | fast 3349.1t/s | fast 1842.0t/s |
| Single GTX 1080 Ti (11GB) | fast 2420.0t/s | fast 1331.0t/s |
| Dual EPYC 9004 + 768GB DDR5-4800 | fast 2304.0t/s | fast 1267.2t/s |
| DGX Spark 128GB unified | fast 1365.0t/s | fast 750.8t/s |
| Ryzen AI Max+ 395 128GB | fast 1280.0t/s | fast 704.0t/s |
| Jetson AGX Orin 64GB | fast 1024.0t/s | fast 563.2t/s |
| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | fast 1024.0t/s | fast 563.2t/s |
| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | fast 768.0t/s | fast 422.4t/s |
| NVIDIA Jetson Orin NX 16GB | fast 512.0t/s | fast 281.6t/s |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options #
Or run it in the cloud #
No per-token API provider pricing tracked for Nemotron 3 Diarization yet.
For flagship list prices, see the
[calculator](https://tokenstead.ai/calculator).
Inference cost over time #
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.