How We Built a Sub-500ms Real-Time AI Voice Agent with TypeScript, Twilio & WebSockets Hikari Webworks detailed a streaming architecture that achieves sub-500ms end-to-end voice latency for AI phone agents by running the entire pipeline over full-duplex WebSockets rather than sequential HTTP request-response cycles. The stack chains Twilio Media Streams, Deepgram Nova-2 streaming speech-to-text, a streaming LLM (Groq Llama 3.3 70B or Claude 3.5 Haiku), and fast TTS (Cartesia Sonic or ElevenLabs Turbo), with per-stage latencies of roughly 80ms, 140ms, 90ms and 40ms for a total turnaround of about 350-480ms. The writeup also describes barge-in handling that dispatches a Twilio 'clear' event and aborts the pending TTS stream within 50ms when a caller interrupts. When building conversational AI phone agents for real-world production environments such as after-hours dental clinic receptionists or real estate lead qualifiers , latency is everything . In human conversation, the natural pause between speakers averages 200ms to 400ms . If an AI telephone system takes 1.5 to 3 seconds to respond: In this tutorial, we share the exact streaming architecture we engineered at Hikari Webworks https://hikariwebworks.studio to achieve sub-500ms end-to-end voice latency in production. Sequential HTTP request-response cycles Record audio - Send POST - Wait for LLM - Generate MP3 - Play audio add at least 2.5 seconds of artificial delay. To achieve sub-500ms, the entire pipeline must operate over full-duplex bi-directional WebSockets : Phone Caller │ 8kHz raw audio stream over PSTN/SIP ▼ Twilio Media Stream WebSocket │ ├── 80ms ──► Deepgram Nova-2 Live Streaming STT │ ├── 140ms ─► Streaming LLM Groq Llama 3.3 70B / Claude 3.5 Haiku │ ├── 90ms ──► Ultra-Fast TTS Cartesia Sonic / ElevenLabs Turbo │ └── 40ms ──► Raw Mu-Law Audio Chunk Streaming back to Caller Total Turnaround Time: ~350ms - 480ms Sub-500ms Real-Time Loop When a call connects, Twilio opens a WebSocket connection and streams 8kHz 16-bit Mu-law audio packets encoded in base64: js // server/voice-agent.ts import { WebSocketServer, WebSocket } from 'ws'; interface TwilioMediaMessage { event: 'connected' | 'start' | 'media' | 'stop' | 'mark' | 'clear'; streamSid?: string; media?: { payload: string; // Base64 encoded 8kHz mu-law audio timestamp: string; chunk: string; }; } export function initializeVoiceServer port: number = 8080 { const wss = new WebSocketServer { port } ; wss.on 'connection', ws: WebSocket = { let streamSid = ''; ws.on 'message', async data: string = { const msg: TwilioMediaMessage = JSON.parse data ; if msg.event === 'start' && msg.streamSid { streamSid = msg.streamSid; console.log Twilio Call stream active: ${streamSid} ; } if msg.event === 'media' && msg.media { // Decode raw audio frame and stream immediately to Speech-to-Text const rawChunk = Buffer.from msg.media.payload, 'base64' ; liveSttStream.send rawChunk ; } if msg.event === 'stop' { console.log Twilio Call ended: ${streamSid} ; } } ; } ; console.log AI Voice Gateway listening on port ${port} ; } One of the most critical engineering challenges in voice AI is handling human interruptions. If the AI is in the middle of speaking and the user says "Wait, how much does that cost?" , the AI must shut up instantly. To achieve this: < 50ms . clear signal is immediately dispatched to Twilio to purge the pending audio playback buffer. // Handling real-time human barge-in function handleUserInterruption ws: WebSocket, streamSid: string { // Purge Twilio's audio queue immediately const clearPayload = JSON.stringify { event: 'clear', streamSid: streamSid } ; ws.send clearPayload ; // Abort pending TTS stream chunk generator ttsStreamController.abort ; } In our Hikari Real Estate AI Solution https://hikariwebworks.studio/solutions/real-estate , the voice agent qualifies buyer budget and timeline, books an executive site tour, and dispatches project brochures over WhatsApp in under 2 seconds: js // Define structured function calling schemas const qualificationTools = { type: "function", function: { name: "confirm site visit and send whatsapp", description: "Confirms property site visit and immediately sends location & e-brochure over WhatsApp.", parameters: { type: "object", properties: { buyerName: { type: "string" }, phoneNumber: { type: "string" }, unitPreference: { type: "string", enum: "2BHK", "3BHK", "Villa" }, budgetLakhs: { type: "number" }, visitDateTime: { type: "string" } }, required: "buyerName", "phoneNumber", "unitPreference", "visitDateTime" } } } ; Across over 50,000 live conversations handled across our AI Phone Agent Deployments https://hikariwebworks.studio/services/phone-agents and Enterprise AI Operations https://hikariwebworks.studio/solutions/ai-operations : To test live interactive ROI benchmarks and voice latency cost estimators, explore our Web & AI Cost Calculator https://hikariwebworks.studio/cost-calculator or connect directly with our engineering team at Hikari Webworks https://hikariwebworks.studio .