Tango: Simple AI's Conversational Awareness Model Simple AI introduced Tango, an audio-native end-of-turn model that in production lands its end-of-turn decision a median 300ms before a conventional voice-activity detector would detect a pause, measured on real telephony audio. On the public LiveKit benchmark, Tango was the fastest of 11 systems at a minimum 90% detection rate, missing 15 of 400 true turn endings versus 54 for Soniox and 197 for Deepgram Flux, and it fires 3.7 to 4.7 times less often mid-speech than Smart Turn at equal turn-end coverage. Simple AI said Tango cut interruptions by 64% versus its previous production agent, recognizing a real interruption about 100ms after the customer starts talking, with assessments made live every 80 ms and inference run in less than half that time. Voice AI’s biggest challenge is timing. For a conversation to feel natural, an agent has to recognize the moment a customer is done speaking and craft their response quickly and accurately. Most models still struggle with this, either interrupting mid-sentence, or pausing too long to confirm their turn to speak. Customers are left with conversations that feel robotic rather than human. It’s the biggest obstacle standing between voice AI and mainstream adoption. We tried all of the end-of-turn models on the market and couldn’t find one that actually worked well in production. In Simple AI fashion, we chose to train our own. Today, we’re introducing Tango , the system behind Simple’s most natural voice agents yet. In production, its end-of-turn decision lands a median 300ms before a conventional voice-activity detector would even detect a pause — measured on real telephony audio, not just clean VoIP. How Voice AI Used to Work Most voice AI on the market uses a cascaded approach: transcribe what the customer says, generate a response from that transcript, convert it back to audio. Everything the agent needs has to survive being flattened into text first, and a lot gets lost in translation. The result is both slower and less natural: no matter how fast your agent is, it’s waiting on that transcript before it can say anything. Tango skips that step for the decisions that matter most, reading end-of-turn and interruption directly from audio. Transcripts still get generated for every call, but they no longer sit in the critical path. Benefits of Tango’s Audio Processing Improved end-of-turn accuracy Tango isn’t just faster than competitor end-of-turn models, it’s also more accurate. This is crucial, since false positives mean more interruptions and false negatives mean uncomfortable pauses. With improved end-of-turn, conversations flow more freely. On the public LiveKit benchmark, it’s the fastest of 11 systems once you require at least 90% detection, missing just 15 of 400 true turn endings versus 54 for Soniox and 197 for Deepgram Flux. Against Smart Turn specifically, it fires 3.7 to 4.7 times less often mid-speech at equal turn-end coverage. Fewer accidental interruptions Genuine interruptions and quick backchannel cues like “mm-hmm” can look identical in a transcript, but they call for opposite behavior: stop talking for one, keep going for the other. Tango tells them apart in real time, so it doesn’t cut off a customer who’s simply agreeing. That’s cut interruptions by 64% versus our previous agent in production, recognizing a real interruption about 100ms after the customer starts talking. Hear what text can’t show Type “thank you” and there’s no way to tell if it was said with a smile or through gritted teeth. Text strips out tone, prosody, and pacing: the exact signals that make an interaction feel human. Audio-native processing keeps that signal intact. Speed, without the translation loss The old pipeline adds roughly one second of latency before your agent starts thinking. Tango skips it: audio comes in, decisions come out, with both agent and customer channels processed at once. By going straight to streaming audio models, we can make assessments live every 80 ms, and run inference in less than half that time. There’s no transcript to wait for, and no second pipeline to run. What This Looks and Sounds Like A transcript-based model sees silence and guesses. Tango hears the customer still thinking, and responds at just the right time. Tango can tell a real interruption from a conversational nod, so your agents don’t stop talking every time a customer says “okay”. What This Means For Your Call Center In practice, the difference between a natural-feeling agent and a robotic one comes down to about one second of latency. This gap shows up directly in the metrics you’re already tracking. - Customer Satisfaction Score CSAT : A customer’s satisfaction is shaped by their overall experience rather than resolution alone. Fewer awkward pauses and false interruptions means fewer calls that feel broken, even when the agent is correct. - Average Handle Time AHT : Every pause, interruption, or repeated “sorry, go ahead” adds seconds to your call. At call center volume, these seconds compound fast. - Abandonment: The customers you lose aren’t the ones who got a wrong answer. They’re the ones who hung up in frustration before you got the chance to help them at all. How We Compare A system near the top of this chart catches more end-of-turns, while a system to the left catches them sooner. Tango is the most accurate model for its speed, and the fastest model for its accuracy. Because the best, most natural conversations have both. Built for the Calls You Actually Get Most end-of-turn models are trained and benchmarked on clean, high-bandwidth audio. Real customer calls are not. They’re compressed, 8kHz telephony audio, running through phone systems built way before there were AI agents on the other end of the line. That’s why models that perform well in a demo often degrade the moment they meet a call center’s actual call volume. Tango is built dual-channel and telephony-native from the ground up. It plugs directly into the inbound and outbound calling systems you already run, without asking you to upgrade your infrastructure first. That’s the difference between a benchmark number and a number you’ll actually see in production. What’s Next for Simple Tango Tango is already live across Simple’s production traffic, delivering lower latency, higher accuracy, and more natural conversations for the calls our customers handle every day. It’s already shaping how we think about every agent we build from here. They say it takes two to Tango. Schedule a demo today to see how Tango could sound for you.