cd /news/artificial-intelligence/nemotron-3-diarization · home › topics › artificial-intelligence › article
[ARTICLE · art-140979] src=tokenstead.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Nemotron 3 Diarization

NVIDIA released Nemotron 3 Diarization, a 100M-parameter open speaker-diarization model that ranks #1 on VoiceArena's Diarization-Bench leaderboard with a 14.72% Diarization Error Rate. The model supports up to eight speakers in real time, cutting DER by an average of 41% relative against NVIDIA's previous four-speaker streaming baseline at 1.04-second latency, and ships under the openmdw-1.1 license for commercial use with GGUF, ONNX, CoreML, and LiteRT community ports. NVIDIA reported 26,000 downloads in five days.

read4 min views1 publishedSep 28, 2026
Nemotron 3 Diarization
Image: Tokenstead (auto-discovered)

edge Who spoke when, in 107 megabytes. Nemotron 3 Diarization is NVIDIA’s open speaker-diarization model: a 100M-parameter Sortformer descendant that decides which of up to eight speakers is talking, in real time, and ranks #1 on VoiceArena’s Diarization-Bench leaderboard at a 14.72% Diarization Error Rate. Against NVIDIA’s previous four-speaker streaming baseline it cuts DER by an average of 41% relative at 1.04-second latency, with the advantage widening as speaker counts rise - the eight-speaker support is what makes it a meetings model rather than a phone-call model.

Streaming and offline in one checkpoint. One checkpoint covers both modes: input buffer as low as 80ms for latency-critical uses (the recommended floor is 0.32s), 30.4s buffer for offline-quality transcripts, output frames in 10ms multiples, and chunked inference with no duration limit. It produces speaker activity and timestamps, not words: chain it with a timestamped ASR model (Nemotron ASR 3.5 or Parakeet TDT 0.6B v3) and word-to-speaker mapping gives you attributed transcripts, the piece every local meeting-assistant pipeline was missing.

What runs where. Two routes. The NeMo path wants Linux with an NVIDIA Ampere-or-newer GPU. The interesting one for this audience is NeMo-Speech.cpp, NVIDIA’s C++ runtime (nemo-speech diarize meeting.wav, or --diarize to tag a transcript with word-level speakers), plus community ports: a GGUF at 198.7 MB bf16 / 106.7 MB q8_0, ONNX, CoreML, and LiteRT builds, which put real-time diarization on CPUs and phones, not just CUDA boxes. License is openmdw-1.1, explicitly ready for commercial use, which is the difference between a demo and a pipeline component. 26,000 downloads in five days say the meeting-tools crowd noticed.

  • 100M
  • openmdw 1.1
  • 🇺🇸 USA
  • Sep 2026

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally #

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

The reference hardware

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

Rig Q8_0 BF16
4x H100 80GB (320GB) fast 67000.0t/s fast 36850.0t/s
NVIDIA DGX Station 748GB fast 40000.0t/s fast 22000.0t/s
8x RTX 3090 rack (192GB) fast 37448.0t/s fast 20596.4t/s
4x RTX 5090 (128GB) fast 35840.0t/s fast 19712.0t/s
AMD Instinct MI300X (192GB) fast 26624.0t/s fast 14643.2t/s
4x RTX 4090 (96GB) fast 20160.0t/s fast 11088.0t/s
2x RTX 5090 (64GB) fast 17920.0t/s fast 9856.0t/s
2x RTX 3090 (48GB) fast 9362.0t/s fast 5149.1t/s
Single RTX 5090 (32GB) fast 8960.0t/s fast 4928.0t/s
RTX PRO 6000 Blackwell (96GB) fast 8960.0t/s fast 4928.0t/s
Mac Studio M4 Ultra 192GB fast 5956.4t/s fast 3276.0t/s
Mac Studio M4 Ultra 512GB fast 5956.4t/s fast 3276.0t/s
Single RTX 4090 (24GB) fast 5040.0t/s fast 2772.0t/s
MacBook Pro M5 Max 128GB fast 3349.1t/s fast 1842.0t/s
Single GTX 1080 Ti (11GB) fast 2420.0t/s fast 1331.0t/s
Dual EPYC 9004 + 768GB DDR5-4800 fast 2304.0t/s fast 1267.2t/s
DGX Spark 128GB unified fast 1365.0t/s fast 750.8t/s
Ryzen AI Max+ 395 128GB fast 1280.0t/s fast 704.0t/s
Jetson AGX Orin 64GB fast 1024.0t/s fast 563.2t/s
Epyc + 512GB DDR4-3200 + 2x RTX 3090 fast 1024.0t/s fast 563.2t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 fast 768.0t/s fast 422.4t/s
NVIDIA Jetson Orin NX 16GB fast 512.0t/s fast 281.6t/s

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options #

Or run it in the cloud #

    No per-token API provider pricing tracked for Nemotron 3 Diarization yet.
        For flagship list prices, see the
        [calculator](https://tokenstead.ai/calculator).

Inference cost over time #

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nemotron-3-diarizati…] indexed:0 read:4min 2026-09-28 · —