Talking to a voice agent feels simple. You speak, it answers. Behind that one exchange, six small models run on every turn, and the hardest question the agent faces is not what to say. It is when to say it.
This article is about that side of the problem. It uses Alex, a voice agent that negotiates freight rates over the phone, as the example. A companion article, How to Stop an AI Agent From Giving Away Your Money in a Negotiation, covers how Alex protects the price. This one covers how Alex listens, answers and keeps a conversation flowing. You can call Alex from a browser at talk-to-alex.web.app.
Much of the vocabulary here comes from DeepLearning.AI’s short course Building AI Voice Agents for Production, taught by the teams at LiveKit and RealAvatar (d’Sa et al., 2025).
People are fast at taking turns. A study of conversations in ten languages found that the gap between one person finishing and the other starting averages about 200 milliseconds (Stivers et al., 2009). We do not wait for silence. We predict the end of a sentence from its words and its tone.
A voice agent cannot be that fast, but it has to get close. On every turn it does four things, the four numbered steps in the diagram above:
Alex is built with LiveKit Agents, an open source framework that wires these steps together. The rest of this article goes through them one by one.
There is a shortcut. Speech to speech models, like OpenAI’s Realtime API or Google’s Gemini Live API, take audio in and give audio out, with no text in between (Google, 2026; OpenAI, n.d.). They are quick and sound natural.
Alex uses a pipeline instead, for one reason: every price it says has to be checked by code before it is spoken, and code can only check text. The pipeline keeps a moment where the answer exists as written words. That moment costs a little time, and most of what follows is about keeping that cost small.
The caller’s voice travels over WebRTC, the same technology behind video calls in a browser. It sends audio in small packets, one every 20 milliseconds, and if one gets lost it simply moves on; the audio codec fills the tiny gap so nobody hears it (Valin & Bran, 2016). That matters more than it sounds. With a regular web connection, one lost packet makes everything behind it wait, and the caller hears a stutter.
On the other side, a speech recognizer (Deepgram Nova-3) writes down what the caller says while they are still saying it. It runs in a multilingual mode, so the same call can switch between English and Spanish (Francisco, 2025). The web page tells Alex which language it is in, and Alex answers in that language.
This was the hardest part of the build, and the one that decides how the call feels.
Two models are involved. The first, voice activity detection, only knows whether there is sound or silence. The second, the turn detector, listens to how the words are said: a voice that drops at the end of a sentence, or stays up because more is coming (LiveKit, n.d.-a).
The difference shows up as soon as someone says a price. People say “I can do it for thirty two…” and before “…fifty”. The first model hears silence and would let Alex answer to $3,200. The turn detector hears a sentence that is not finished and waits for the real number, $3,250.
The turn detector gives a probability, not a yes or no, and two settings decide how long Alex waits:
The first version felt slow. The measurements showed why: whenever the detector was unsure, the framework waited its default three seconds before answering (LiveKit, n.d.-b). That is a safe default for a general assistant. On a sales call, three seconds of silence sounds like a dropped line. Alex now waits 0.3 seconds when the detector is sure and at most 1.2 seconds when it is not, which is still long enough for a in the middle of a number. In code, it is two numbers:
turn_handling=TurnHandlingOptions( turn_detection=inference.TurnDetector(), endpointing=EndpointingOptions(min_delay=0.3, max_delay=1.2),)
For a voice agent, the most important property of a language model is not how smart it is. It is how quickly it starts answering, because everything before the first word is silence on the line.
Alex’s first model was chosen for cost. The first conversation took three to four seconds per turn, most of it waiting for the model. So three models were timed on the same short prompts, through the same service:
GPT-4.1 mini started answering in about 0.7 seconds, three times faster than the first choice. The test took two minutes and changed the decision. Two rules came with it: no models that “think” before answering, since thinking is dead air, and a low temperature, because a negotiator should be consistent.
The model chooses the words, never the numbers. How that works, and what it costs, is the subject of the companion article.
The voice (Cartesia Sonic) starts speaking as soon as the first sentence is ready, while the model is still writing the second one (Cartesia, n.d.). The caller hears the start of the answer before the end of it exists.
Real callers also talk over the agent. They say “yeah, yeah” or cut in with a counter offer.
When that happens, Alex stops talking right away. Then the framework does something small and important: it trims Alex’s memory to the part the caller actually heard (LiveKit, n.d.-c). If Alex was cut off after “The load picks up Monday”, the model does not remember mentioning Dallas, so it never assumes the caller knows.
Alex also hangs up on its own. When the deal is done or the caller says goodbye, it waits for its last sentence to finish playing and then closes the call.
Every turn is timed, from the caller’s last word to Alex’s first sound.
Across every call answered so far, the median turn took 1.39 seconds, and nine in ten took less than 1.81. The breakdown of one full call shows where the time goes: about 1.2 seconds waiting to be sure the caller had finished, about 0.4 seconds for the model to start, about 0.1 seconds for the voice.
The biggest block is not the AI. It is waiting for the turn to end. That is the trade at the heart of every voice agent: answer sooner and you will sometimes interrupt, wait longer and every turn feels slow.
Alex runs on Google Cloud as two services: the web page and the agent.
They behave in opposite ways. The web page only works when someone visits, so it can switch off when nobody is there. The agent cannot. It keeps a line open and waits for calls, and if the platform s it to save resources, the next call rings with nobody to answer. So the agent runs on one instance that is always on with its processor always active (Google Cloud, 2026). It is the only fixed cost of the project.
The first deployments also taught three lessons worth passing on:
The latency numbers come from a few dozen turns, enough to see where the time goes but not enough to promise a number. Callers reach Alex from a browser, not a phone line, and the waiting times were tuned by hand in quiet rooms. Accents, background noise and misheard numbers have not been tested at scale.
The next step is a test bench: a second voice agent that plays the caller and calls Alex hundreds of times, with noise, accents and mumbled numbers, to measure how often Alex interrupts or mishears a price.
Most of the work in a voice agent is not the language model. Choosing the model took two minutes once it was measured. The real work is timing: moving audio without stalls, knowing when someone has finished, starting to speak before the answer is complete, stopping the moment the caller talks, and remembering only what was actually said.
The code and every measurement are on GitHub at voice-freight-negotiator (Peña Donneys, 2026), and Alex answers calls at talk-to-alex.web.app.
Cartesia. (n.d.). Sonic: The fastest and most natural text to speech model. Retrieved October 3, 2026, from https://www.cartesia.ai/sonic
d’Sa, R., Parmelee, S., & Teneva, N. (2025). Building AI voice agents for production [Online course]. DeepLearning.AI. https://www.deeplearning.ai/courses/building-ai-voice-agents-for-production
Francisco, J. N. (2025, February 12). Introducing Nova-3: Setting a new standard for AI-driven speech-to-text. Deepgram. https://deepgram.com/learn/introducing-nova-3-speech-to-text-api
Google. (2026, September 15). Gemini Live API overview. Google AI for Developers. https://ai.google.dev/gemini-api/docs/live-api
Google Cloud. (2026, September 30). Billing settings for services. Cloud Run documentation. https://docs.cloud.google.com/run/docs/configuring/billing-settings
LiveKit. (n.d.-a). LiveKit turn detector. LiveKit Documentation. Retrieved October 3, 2026, from https://docs.livekit.io/agents/logic/turns/turn-detector/
LiveKit. (n.d.-b). Turn handling options. LiveKit Documentation. Retrieved October 3, 2026, from https://docs.livekit.io/reference/agents/turn-handling-options/
LiveKit. (n.d.-c). Turns overview. LiveKit Documentation. Retrieved October 3, 2026, from https://docs.livekit.io/agents/logic/turns/
OpenAI. (n.d.). Getting started with the Realtime API. OpenAI API. Retrieved October 3, 2026, from https://developers.openai.com/api/docs/guides/realtime
Peña Donneys, J. S. (2026). Voice freight negotiator [Computer software]. GitHub. https://github.com/JSebastianIEU/voice-freight-negotiator
Stivers, T., Enfield, N. J., Brown, P., Englert, C., Hayashi, M., Heinemann, T., Hoymann, G., Rossano, F., de Ruiter, J. P., Yoon, K.-E., & Levinson, S. C. (2009). Universals and cultural variation in turn-taking in conversation. Proceedings of the National Academy of Sciences, 106(26), 10587–10592. https://doi.org/10.1073/pnas.0903616106
Valin, J.-M., & Bran, C. (2016). WebRTC audio codec and processing requirements (RFC 7874). RFC Editor. https://doi.org/10.17487/RFC7874
All images are by the author, Juan Sebastian Peña Donneys, designed in Figma.
What It Takes to Build a Voice Agent That Can Hold a Phone Call was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.