# PolyAI launches new real-time voice conversation model to make AI-driven calls more human

> Source: <https://siliconangle.com/2026/07/30/polyai-launches-new-real-time-voice-conversation-model-make-ai-driven-calls-human/>
> Published: 2026-07-30 13:00:48+00:00

### PolyAI launches new real-time voice conversation model to make AI-driven calls more human

Voice assistant and conversational AI agent developer [PolyAI Ltd.](https://poly.ai/) today announced the release of Dialog-RSN-1, a voice dialog artificial intelligence model capable of directly perceiving and responding to audio.

Voice AI companies have been working to reduce delay and increase accuracy in voice AI agents, enabling them to sound more human on phone calls. In many cases, the chain of software needed to make this happen usually starts with speech recognition and processing, then the large language model for text generation, and finally speech output.

The company’s new LLM does speech recognition and processing all within the model, but speech output is split out. This allows the AI model to completely “hear” and understand the voice conversation all at once, without another model in front of it that cannot pick up tone, cadence, or understand the full context of the conversation.

According to PolyAI, this new model provides an experience closer to another human listening to the call. Having the audio directly parsed by the AI, it can pick up on non-speech cues that might change the path of the conversation.

This also allows users to prompt the model to react to what it hears in a caller’s accent, mood or noise in the environment. Instead of working from a transcript, which loses all of this, the model can catch mispronounced words made obvious by context, a caller half-spelling out loud “Matthew with two Ts” or a brand name like Audibel, which isn’t the word “audible.”

A benefit of this is that Dialog-RSN can react quickly and catch potential delays on the line. This means that the model can respond quickly, run tools under the hood, and build based on the entire context of the conversation. It also reduces the likelihood the model will interrupt the user, because it can catch onto the cadence and timing of the user’s speech, giving it a similar insight into how normal humans take turns when conversing.

According to the company, Dialog-RSN has a latency between 280 and 500 milliseconds, meaning it reliably responds in under 300ms. Comparatively, OpenAI GPT’s GPT-realtime-2.1 often responds in between 860 and 1900ms.

The natural delay between speakers in a normal human conversation is around 200 – 300ms, or around 2 tenths of a second. This delay is almost universal and remarkably consistent across different languages, cultures and geographic regions. Humans catch onto deviations from conversational rhythm; that timing is important not just for conversation but for creating an understanding. Slow responses from an AI agent can make it feel frustrating or awkward. Although this is not always that big of a deal with automated systems – most people don’t expect a machine on the line to respond like a human – it leads to a problem when a caller wants to interrupt.

If an AI agent is going on about something the user doesn’t feel like it caught correctly (or even another human), most people will interrupt so that it doesn’t take another minute or more to keep going on a conversation turn that’s mistaken or taking up time when they could clarify or move on.

The company also focused on splitting out the text-to-speech module from the LLM itself; this allows users to more readily control the weights and output. Having Dialog-RSN provide its outputs to a TTS model means that the end user can easily tune their speech outputs with a voice, emotionality and cadence that they want without needing to adjust the conversational model – which would require additional and costly training.

According to PolyAI, keeping generation separate gives full control over the output voice. This means that switching out a voice, changing its intonation, accent, emotional resonance and other customizations can happen without adjusting the underlying model.

The comparison is to voice models such as GPT realtime and Gemini Live, where the output voice is baked into the model itself. This makes alternate voices difficult to control at scale and more costly to stream. Since the TTS is separate, companies also control how much spend they want to run through speech generation, run it on its own hardware and make it as efficient as possible, thus reducing the overall compute bill.

PolyAI said it built the model through post-training of open-weight multimodal models using supervised fine-tuning and reinforcement tuning using in-house training data. The company said its training pipeline is broadly model-agnostic, so it evaluated Gemma, GPT-OSS, Qwen and Mistral. The aim was to hit the sub-300ms delay when served on A100 graphics processing units using an 8B dense or 30B sparse model.

##### Image: SiliconANGLE/Microsoft Designer

# A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. **Join theCUBE’s Alumni Trust Network**, where technology leaders connect, share intelligence and create opportunities.

**15M+ viewers of theCUBE videos**, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.

# Are you AWS customer? Support SiliconANGLE Financially by buying your AWS services from our Marketplace portal page and links.

**About SiliconANGLE Media**

[SiliconANGLE](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552),

[theCUBE Network](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da),

[theCUBE Research](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f),

[CUBE365](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6),

[theCUBE AI](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683)and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.
