{"slug": "how-we-built-a-sub-500ms-real-time-ai-voice-agent-with-typescript-twilio", "title": "How We Built a Sub-500ms Real-Time AI Voice Agent with TypeScript, Twilio & WebSockets", "summary": "Hikari Webworks detailed a streaming architecture that achieves sub-500ms end-to-end voice latency for AI phone agents by running the entire pipeline over full-duplex WebSockets rather than sequential HTTP request-response cycles. The stack chains Twilio Media Streams, Deepgram Nova-2 streaming speech-to-text, a streaming LLM (Groq Llama 3.3 70B or Claude 3.5 Haiku), and fast TTS (Cartesia Sonic or ElevenLabs Turbo), with per-stage latencies of roughly 80ms, 140ms, 90ms and 40ms for a total turnaround of about 350-480ms. The writeup also describes barge-in handling that dispatches a Twilio 'clear' event and aborts the pending TTS stream within 50ms when a caller interrupts.", "body_md": "When building conversational AI phone agents for real-world production environments (such as after-hours dental clinic receptionists or real estate lead qualifiers), **latency is everything**.\n\nIn human conversation, the natural pause between speakers averages **200ms to 400ms**. If an AI telephone system takes 1.5 to 3 seconds to respond:\n\nIn this tutorial, we share the exact streaming architecture we engineered at [Hikari Webworks](https://hikariwebworks.studio) to achieve **sub-500ms end-to-end voice latency** in production.\n\nSequential HTTP request-response cycles (`Record audio -> Send POST -> Wait for LLM -> Generate MP3 -> Play audio`) add at least 2.5 seconds of artificial delay.\n\nTo achieve sub-500ms, the entire pipeline must operate over **full-duplex bi-directional WebSockets**:\n\n```\n[Phone Caller] \n      │ (8kHz raw audio stream over PSTN/SIP)\n      ▼\n[Twilio Media Stream WebSocket]\n      │\n      ├── (80ms) ──► Deepgram Nova-2 (Live Streaming STT)\n      │\n      ├── (140ms) ─► Streaming LLM (Groq Llama 3.3 70B / Claude 3.5 Haiku)\n      │\n      ├── (90ms) ──► Ultra-Fast TTS (Cartesia Sonic / ElevenLabs Turbo)\n      │\n      └── (40ms) ──► Raw Mu-Law Audio Chunk Streaming back to Caller\n\nTotal Turnaround Time: ~350ms - 480ms (Sub-500ms Real-Time Loop)\n```\n\nWhen a call connects, Twilio opens a WebSocket connection and streams 8kHz 16-bit Mu-law audio packets encoded in base64:\n\n``` js\n// server/voice-agent.ts\nimport { WebSocketServer, WebSocket } from 'ws';\n\ninterface TwilioMediaMessage {\n  event: 'connected' | 'start' | 'media' | 'stop' | 'mark' | 'clear';\n  streamSid?: string;\n  media?: {\n    payload: string; // Base64 encoded 8kHz mu-law audio\n    timestamp: string;\n    chunk: string;\n  };\n}\n\nexport function initializeVoiceServer(port: number = 8080) {\n  const wss = new WebSocketServer({ port });\n\n  wss.on('connection', (ws: WebSocket) => {\n    let streamSid = '';\n\n    ws.on('message', async (data: string) => {\n      const msg: TwilioMediaMessage = JSON.parse(data);\n\n      if (msg.event === 'start' && msg.streamSid) {\n        streamSid = msg.streamSid;\n        console.log(`[Twilio] Call stream active: ${streamSid}`);\n      }\n\n      if (msg.event === 'media' && msg.media) {\n        // Decode raw audio frame and stream immediately to Speech-to-Text\n        const rawChunk = Buffer.from(msg.media.payload, 'base64');\n        liveSttStream.send(rawChunk);\n      }\n\n      if (msg.event === 'stop') {\n        console.log(`[Twilio] Call ended: ${streamSid}`);\n      }\n    });\n  });\n\n  console.log(`AI Voice Gateway listening on port ${port}`);\n}\n```\n\nOne of the most critical engineering challenges in voice AI is handling human interruptions. If the AI is in the middle of speaking and the user says *\"Wait, how much does that cost?\"*, the AI must shut up instantly.\n\nTo achieve this:\n\n`< 50ms`.` clear` signal is immediately dispatched to Twilio to purge the pending audio playback buffer.\n\n```\n// Handling real-time human barge-in\nfunction handleUserInterruption(ws: WebSocket, streamSid: string) {\n  // Purge Twilio's audio queue immediately\n  const clearPayload = JSON.stringify({\n    event: 'clear',\n    streamSid: streamSid\n  });\n\n  ws.send(clearPayload);\n\n  // Abort pending TTS stream chunk generator\n  ttsStreamController.abort();\n}\n```\n\nIn our [Hikari Real Estate AI Solution](https://hikariwebworks.studio/solutions/real-estate), the voice agent qualifies buyer budget and timeline, books an executive site tour, and dispatches project brochures over WhatsApp in under 2 seconds:\n\n``` js\n// Define structured function calling schemas\nconst qualificationTools = [\n  {\n    type: \"function\",\n    function: {\n      name: \"confirm_site_visit_and_send_whatsapp\",\n      description: \"Confirms property site visit and immediately sends location & e-brochure over WhatsApp.\",\n      parameters: {\n        type: \"object\",\n        properties: {\n          buyerName: { type: \"string\" },\n          phoneNumber: { type: \"string\" },\n          unitPreference: { type: \"string\", enum: [\"2BHK\", \"3BHK\", \"Villa\"] },\n          budgetLakhs: { type: \"number\" },\n          visitDateTime: { type: \"string\" }\n        },\n        required: [\"buyerName\", \"phoneNumber\", \"unitPreference\", \"visitDateTime\"]\n      }\n    }\n  }\n];\n```\n\nAcross over 50,000 live conversations handled across our [AI Phone Agent Deployments](https://hikariwebworks.studio/services/phone-agents) and [Enterprise AI Operations](https://hikariwebworks.studio/solutions/ai-operations):\n\nTo test live interactive ROI benchmarks and voice latency cost estimators, explore our [Web & AI Cost Calculator](https://hikariwebworks.studio/cost-calculator) or connect directly with our engineering team at [Hikari Webworks](https://hikariwebworks.studio).", "url": "https://wpnews.pro/news/how-we-built-a-sub-500ms-real-time-ai-voice-agent-with-typescript-twilio", "canonical_source": "https://dev.to/devtrivedi/how-we-built-a-sub-500ms-real-time-ai-voice-agent-with-typescript-twilio-websockets-3hlg", "published_at": "2026-10-10 05:30:00+00:00", "updated_at": "2026-10-10 05:30:21.940930+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "natural-language-processing", "developer-tools"], "entities": ["Hikari Webworks", "Twilio", "Deepgram", "Groq", "Llama 3.3 70B", "Claude 3.5 Haiku", "Cartesia Sonic", "ElevenLabs"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-we-built-a-sub-500ms-real-time-ai-voice-agent-with-typescript-twilio", "markdown": "https://wpnews.pro/news/how-we-built-a-sub-500ms-real-time-ai-voice-agent-with-typescript-twilio.md", "text": "https://wpnews.pro/news/how-we-built-a-sub-500ms-real-time-ai-voice-agent-with-typescript-twilio.txt", "jsonld": "https://wpnews.pro/news/how-we-built-a-sub-500ms-real-time-ai-voice-agent-with-typescript-twilio.jsonld"}}