cd /news/artificial-intelligence/nvidia-drops-a-free-100m-parameter-m… · home › topics › artificial-intelligence › article
[ARTICLE · art-140435] src=the-decoder.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Nvidia drops a free 100M-parameter model that identifies up to eight speakers in real time

Nvidia released Nemotron 3 Diarization, a roughly 100-million-parameter AI model that identifies which speaker is talking at any moment and can distinguish up to eight speakers, with weights freely available on Hugging Face. On the Diarization-Bench from VoiceArena, the model ranks first with a 14.72 percent error rate, ahead of the next best system at 19.3 percent, and cuts error by an average of 41 percent across eight test scenarios versus its predecessor Streaming Sortformer using a 1.04-second buffer. The model works with both recordings and live audio and, paired with a speech recognition system like Parakeet, produces transcripts with anonymous speaker labels such as "speaker_2.

by read1 min views1 publishedSep 27, 2026
Nvidia drops a free 100M-parameter model that identifies up to eight speakers in real time
Image: The Decoder

Nvidia released Nemotron 3 Diarization, an AI model that identifies which speaker is talking at any given moment in a conversation. The model has about 100 million parameters, and its weights are freely available. It can tell apart up to eight speakers and detect when multiple people talk at the same time. More participants, heavy background noise, or reverb push error rates higher. Paired with a speech recognition system like Parakeet, the model can produce transcripts with speaker labels, though only anonymous ones like "speaker_2." It works with both recordings and live audio.

The audio buffer can be set to four levels ranging from 30.4 down to 0.32 seconds. Shorter buffers generally reduce accuracy. On the Diarization-Bench from VoiceArena, the model currently sits in first place with a 14.72 percent error rate, ahead of the next best system at 19.3 percent. The benchmark is strict. Overlapping speech counts, and even tiny misalignments at speaker transitions are scored as errors. Compared to its predecessor, Streaming Sortformer, the new model cuts the error rate by an average of 41 percent across eight test scenarios when using a 1.04-second buffer.

AI News Without the Hype – Curated by Humans

					Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.				

					Subscribe now

Hugging Face

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nvidia-drops-a-free-…] indexed:0 read:1min 2026-09-27 · —