{"slug": "nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech", "title": "NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling", "summary": "NVIDIA released NemotronLabs VoiceChat 11B, an open 11B end-to-end speech-to-speech model for real-time full-duplex conversation, achieving 448 ms smooth turn-taking latency on Full-Duplex-Bench 1.0 and a user-interruption take-over rate of 1.00 at 480 ms. The model, built from a Fast Conformer encoder, Nemotron Nano v2 LLM, and NVIDIA TTS decoder, supports live tool calling via a separate output channel with on-hold messages, but NVIDIA states it is 'ready for research purposes only' and documents failure modes including a two-minute audio context ceiling and degradation into gibberish after several turns.", "body_md": "NVIDIA has released [NemotronLabs VoiceChat 11B](https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B), an open 11B end-to-end speech-to-speech model for real-time, full-duplex conversation. Instead of chaining ASR, an LLM, and TTS, it performs streaming speech understanding and speech generation in one unified network. That removes the multi-model orchestration and API handoffs a cascaded stack requires, and cuts end-to-end latency: measured smooth turn-taking latency is 448 ms on [Full-Duplex-Bench 1.0](https://arxiv.org/abs/2503.04721). The model listens while it speaks, so a user can barge in mid-turn and the agent yields, with a take-over rate of 1.00 at 480 ms. It is also first open full-duplex model to support tool calling while conversation keeps flowing, using a separate output channel for `<TOOLCALL>`\n\nscripts along with operator-defined “on-hold” lines that fill the gap while an API runs.\n\n**Is it deployable?**\n\n**PARTIAL — deployable today for pilots, not for production.** Weights and container are both public, and the license is permissive. But NVIDIA team states the checkpoint is ‘ready for research purposes only,’ and the repo documents real failure modes: a two-minute audio context ceiling, degradation into non-recoverable gibberish after several turns, runaway self-talk after a turn ends, and dropped words in user transcription.\n\n**Which companies:** any team that can allocate one GPU with at least 80 GB of VRAM — A100, H100, RTX 6000 Pro, or B200 on x86_64 Linux. That covers AI-native startups, funded scaleups, enterprise R&D and innovation labs, GPU cloud providers, and university speech groups. There is no hosted API and no inference provider currently serves the model, so teams without GPU access may not evaluate it.**Industries:** contact centers and CX platforms, automotive in-cabin assistants, retail and drive-thru ordering, telecom IVR modernization, games and NPC dialogue, and accessibility tooling.**Applications:** barge-in-capable voice agents, voice front-ends over internal APIs, live-lookup assistants (weather, pricing, order status), and duplex latency benchmarking harnesses.\n\n**Architecture**\n\n**The model is a hybrid Mamba/Transformer, assembled from three existing NVIDIA components along with one new output path:**\n\n- A\n**Fast Conformer speech encoder** from[Nemotron-Speech-Streaming-En-0.6b](https://huggingface.co/nvidia/nemotron-speech-streaming-en-0.6b), which encodes the incoming 16 kHz stream continuously. - The\n, which consumes audio tokens and predicts text tokens.[NVIDIA Nemotron Nano v2](https://huggingface.co/nvidia/NVIDIA-Nemotron-Nano-9B-v2)LLM backbone - An\n**NVIDIA TTS decoder and codec** that predicts audio codes, rendered as 22.05 kHz agent speech. - A\n**separate output channel** dedicated to tool-calling scripts.\n\nOutputs include agent audio, agent text, and a running user transcription. Training used roughly 550k hours of audio across real and synthetic corpora, building on [SALM-Duplex](https://arxiv.org/abs/2505.15670) and [Audio Flamingo 3](https://arxiv.org/abs/2507.08128).\n\n**Tool calling without dead air**\n\nTool calls are emitted on the side channel as a `<TOOLCALL>`\n\nblock; your code returns results in a `<TOOL_RESPONSE>`\n\nblock. The notable piece is the **on-hold message**: per tool, an operator defines a line the agent speaks the moment the model generates the text triggering the call, so the conversation does not fall silent while an API runs.\n\nConstraints are explicit. NVIDIA recommends a maximum of five tools per session, the model cannot reliably call multiple tools simultaneously, and the user cannot interrupt the agent during tool execution. System prompts and tool responses must be ASCII-only and TTS-friendly.\n\n**Performance**\n\nOn [Full-Duplex-Bench 1.0](https://arxiv.org/abs/2503.04721): smooth turn-taking TOR 0.82 at 448 ms, user-interruption TOR 1.00 at 480 ms, and pause-handling TOR of 0.153 (synthetic) and 0.255 (Candor), where lower is better.\n\nOn [AU Harness](https://github.com/ServiceNow/AU-Harness) BFCL-v3 spoken tool calling: 58.5% simple, 62.5% multiple, 42.5% parallel, 27.5% parallel-multiple, 89.6% irrelevance, **56.1% average**. On [Full-Duplex-Bench v3](https://arxiv.org/abs/2604.04847): 82.5% tool selection, 44.2% argument accuracy, 33% pass@1.\n\nNVIDIA reports the model ranks **#2 among open full-duplex models** on [VoiceBench](https://arxiv.org/abs/2410.17196) and #2 among open models on Full-Duplex-Bench 1.0.\n\n**Interactive explainer**\n\n**Key Takeaways**\n\n- One 11B model replaces the ASR → LLM → TTS chain, at 448 ms measured turn-taking latency.\n- First open full-duplex model with tool calling, using a side channel plus operator-defined on-hold messages.\n- Weights are OpenMDW-1.1 permissive, but NVIDIA labels the checkpoint research-only.\n- Requires one 80 GB GPU; no hosted API exists today.\n\n**Check out the Hugging Face model card**,\n\n[and](https://github.com/NVIDIA-NeMo/Speech/tree/nemotron-labs-voicechat)\n\n**GitHub (NeMo Speech)**[.](https://catalog.ngc.nvidia.com/orgs/nim/nvidia/containers/nemotron-labs-voicechat)\n\n**NGC container****Also, feel free to follow us on**\n\n**and don’t forget to join our**[Twitter](https://x.com/intent/follow?screen_name=marktechpost)\n\n**and Subscribe to**\n\n[150k+ML SubReddit](https://www.reddit.com/r/machinelearningnews/)**. Wait! are you on telegram?**\n\n[our Newsletter](https://www.aidevsignals.com/)\n\n[now you can join us on telegram as well.](https://t.me/machinelearningresearchnews)Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? [Connect with us](https://forms.gle/wbash1wF6efRj8G58)\n\nAsif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.", "url": "https://wpnews.pro/news/nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech", "canonical_source": "https://www.marktechpost.com/2026/08/09/nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech-model-with-450-ms-turn-taking-and-live-tool-calling/", "published_at": "2026-08-09 23:58:34+00:00", "updated_at": "2026-08-10 00:16:09.479191+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "generative-ai", "ai-products"], "entities": ["NVIDIA", "NemotronLabs VoiceChat 11B", "Full-Duplex-Bench 1.0", "Nemotron-Speech-Streaming-En-0.6b", "NVIDIA Nemotron Nano v2", "SALM-Duplex", "Audio Flamingo 3"], "alternates": {"html": "https://wpnews.pro/news/nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech", "markdown": "https://wpnews.pro/news/nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech.md", "text": "https://wpnews.pro/news/nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech.txt", "jsonld": "https://wpnews.pro/news/nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech.jsonld"}}