Nvidia Open Sources a Voice AI That Talks and Listens at Once Nvidia released PersonaPlex-7B-v1 on January 15, 2026, a 7-billion-parameter full-duplex voice model with weights on Hugging Face and code on GitHub under an MIT license and the NVIDIA Open Model License, running on a single GPU. Nvidia reports a 90.8% turn-taking rate, interruption response latency as low as 240 milliseconds, and time-to-first-token of roughly 170 milliseconds, built on the dual-stream Transformer architecture from Kyutai's Moshi. The release matters because full-duplex speech models have largely been locked behind paid APIs, and PersonaPlex ships as a free open model with voice and text persona prompts. Nvidia released PersonaPlex-7B, a free, open-source voice model that scraps the biggest annoyance in talking to a machine: waiting for it to finish before you can speak. Every voice assistant you've used runs on turns. You talk, it waits, it replies, you wait. That rhythm feels fine for a search query. It falls apart the second a conversation gets real, because real conversation is messy. People interrupt. They say "mm-hmm" while someone else is mid-sentence. They talk over each other and sort it out in half a second. Nvidia's new model, PersonaPlex-7B, is built to do that instead of the stiff call-and-response most voice AI still does. Nvidia released PersonaPlex-7B-v1 on January 15, 2026, publishing the weights on Hugging Face and the code on GitHub the same day. It's a 7-billion-parameter model, and it's genuinely free: the code carries an MIT license, the weights run under the NVIDIA Open Model License, and the whole thing runs on a single GPU rather than a server farm. The standard approach to voice AI is really three separate systems stapled together: automatic speech recognition turns your voice into text, a language model writes a reply, and a text-to-speech engine reads that reply out loud. Each handoff adds delay, and those delays stack. The result is the familiar beat of silence before a smart speaker answers you, often north of a second, and a hard rule that only one party can be "active" at a time. PersonaPlex throws that pipeline out. It's full-duplex, meaning it listens and speaks on the same model at the same time, the way a phone call works rather than the way a walkie-talkie does. Architecturally, it does this with a dual-stream design: one stream continuously processes the audio coming in from the user, the other continuously generates the agent's speech and text, and both streams share the same underlying model state. Nothing has to hand off control to something else, because nothing was ever fully off. What Beam Is and Why Nvidia Is Betting Big on Reflection AI https://startupfortune.com/what-beam-is-and-why-nvidia-is-betting-big-on-reflection-ai/ Reflection AI's 501-billion-parameter model trails DeepSeek and Kimi K3 on several benchmarks it published itself, yet Nvidia has backed the startup with $800 million and chip supply deals worth billions through 2029. - why Nvidia backed Reflection AI's Beam model https://startupfortune.com/what-beam-is-and-why-nvidia-is-betting-big-on-reflection-ai/ - Beam 501 billion parameter open source AI model https://startupfortune.com/what-beam-is-and-why-nvidia-is-betting-big-on-reflection-ai/ The model isn't built from scratch. It's based on the dual-stream Transformer architecture from Moshi, the open conversational model built by the French lab Kyutai, with Nvidia fine-tuning and extending it into PersonaPlex. That lineage matters: Nvidia didn't invent full-duplex speech AI, it took an existing open approach and pushed it to a point where the numbers actually hold up. The company reports a 90.8% turn-taking rate, interruption response latency as low as 240 milliseconds, and a time-to-first-token of roughly 170 milliseconds. Put plainly, if you cut the model off mid-sentence, it notices and reacts in roughly a quarter of a second, not the multi-second stall you get from a cascaded system re-running its whole pipeline. There's a control layer on top of the speech engine, too. Before a conversation starts, PersonaPlex takes two prompts: a voice prompt built from audio tokens that sets vocal characteristics and speaking style, and a text prompt that defines the persona itself, its role, background, and the scenario it's operating in. That's what lets the same underlying model show up as a brisk customer service rep in one deployment and a patient tutor in another, without retraining anything. Frankly, the open-source part is the part that should worry competitors more than the latency numbers. Full-duplex speech models have mostly lived behind API paywalls or inside demo videos. Nvidia just put a working one, trained and benchmarked, into anyone's hands for nothing, running on hardware a single developer can own. That's the same playbook chip and software companies have run before: commoditize the layer just below your actual product so more people build on top of it, and more of them end up buying Nvidia GPUs to run it well. It also sets a public benchmark the rest of the industry now has to answer. A 240-millisecond interruption response and a 90.8% turn-taking rate are not vague marketing claims, they're numbers a developer can reproduce locally and hold competing models against. That's a different kind of pressure than another splashy launch video. Voice AI has spent years promising it would feel natural any year now. PersonaPlex is the first time that promise shipped with open weights, a GitHub repo, and a number attached to how often it actually gets the timing right. Also read: AI borrowers now dominate half of America's investment-grade bond market https://startupfortune.com/ai-borrowers-now-dominate-half-of-americas-investment-grade-bond-market/ • AI reverse engineering turns GTA 5 and Halo into browser games overnight https://startupfortune.com/ai-reverse-engineering-turns-gta-5-and-halo-into-browser-games-overnight/ • Elon Musk Told Users to Photograph Their Credit Cards for Grok Bot and His Own Platform Roasted Him for It https://startupfortune.com/elon-musk-told-users-to-photograph-their-credit-cards-for-grok-bot-and-his-own-platform-roasted-him-for-it/ This article is posted in AI News https://startupfortune.com/category/ai/ , check it out for more related stories. Nvidia is in talks to buy the open-source AI lab it just backed https://startupfortune.com/nvidia-is-in-talks-to-buy-the-open-source-ai-lab-it-just-backed/ Reflection AI's valuation has tripled to $25 billion in under a year, and Nvidia, already an $800 million investor, is weighing options from an acqui-hire to a bigger equity stake, Bloomberg and the Financial Times report. - Nvidia acquisition talks with Reflection AI lab https://startupfortune.com/nvidia-is-in-talks-to-buy-the-open-source-ai-lab-it-just-backed/ - how Nvidia is buying open source AI startups https://startupfortune.com/nvidia-is-in-talks-to-buy-the-open-source-ai-lab-it-just-backed/ Join the discussion Open in the community → https://startupfortune.com/community/ Almost there. Sign in and your reply posts straight away.