cd /news/ai-agents/how-we-built-a-sub-500ms-real-time-a… · home › topics › ai-agents › article
[ARTICLE · art-148619] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

How We Built a Sub-500ms Real-Time AI Voice Agent with TypeScript, Twilio & WebSockets

Hikari Webworks detailed a streaming architecture that achieves sub-500ms end-to-end voice latency for AI phone agents by running the entire pipeline over full-duplex WebSockets rather than sequential HTTP request-response cycles. The stack chains Twilio Media Streams, Deepgram Nova-2 streaming speech-to-text, a streaming LLM (Groq Llama 3.3 70B or Claude 3.5 Haiku), and fast TTS (Cartesia Sonic or ElevenLabs Turbo), with per-stage latencies of roughly 80ms, 140ms, 90ms and 40ms for a total turnaround of about 350-480ms. The writeup also describes barge-in handling that dispatches a Twilio 'clear' event and aborts the pending TTS stream within 50ms when a caller interrupts.

by read3 min views1 publishedOct 10, 2026

When building conversational AI phone agents for real-world production environments (such as after-hours dental clinic receptionists or real estate lead qualifiers), latency is everything.

In human conversation, the natural between speakers averages 200ms to 400ms. If an AI telephone system takes 1.5 to 3 seconds to respond:

In this tutorial, we share the exact streaming architecture we engineered at Hikari Webworks to achieve sub-500ms end-to-end voice latency in production.

Sequential HTTP request-response cycles (Record audio -> Send POST -> Wait for LLM -> Generate MP3 -> Play audio) add at least 2.5 seconds of artificial delay.

To achieve sub-500ms, the entire pipeline must operate over full-duplex bi-directional WebSockets:

[Phone Caller] 
      │ (8kHz raw audio stream over PSTN/SIP)
      ▼
[Twilio Media Stream WebSocket]
      │
      ├── (80ms) ──► Deepgram Nova-2 (Live Streaming STT)
      │
      ├── (140ms) ─► Streaming LLM (Groq Llama 3.3 70B / Claude 3.5 Haiku)
      │
      ├── (90ms) ──► Ultra-Fast TTS (Cartesia Sonic / ElevenLabs Turbo)
      │
      └── (40ms) ──► Raw Mu-Law Audio Chunk Streaming back to Caller

Total Turnaround Time: ~350ms - 480ms (Sub-500ms Real-Time Loop)

When a call connects, Twilio opens a WebSocket connection and streams 8kHz 16-bit Mu-law audio packets encoded in base64:

// server/voice-agent.ts
import { WebSocketServer, WebSocket } from 'ws';

interface TwilioMediaMessage {
  event: 'connected' | 'start' | 'media' | 'stop' | 'mark' | 'clear';
  streamSid?: string;
  media?: {
    payload: string; // Base64 encoded 8kHz mu-law audio
    timestamp: string;
    chunk: string;
  };
}

export function initializeVoiceServer(port: number = 8080) {
  const wss = new WebSocketServer({ port });

  wss.on('connection', (ws: WebSocket) => {
    let streamSid = '';

    ws.on('message', async (data: string) => {
      const msg: TwilioMediaMessage = JSON.parse(data);

      if (msg.event === 'start' && msg.streamSid) {
        streamSid = msg.streamSid;
        console.log(`[Twilio] Call stream active: ${streamSid}`);
      }

      if (msg.event === 'media' && msg.media) {
        // Decode raw audio frame and stream immediately to Speech-to-Text
        const rawChunk = Buffer.from(msg.media.payload, 'base64');
        liveSttStream.send(rawChunk);
      }

      if (msg.event === 'stop') {
        console.log(`[Twilio] Call ended: ${streamSid}`);
      }
    });
  });

  console.log(`AI Voice Gateway listening on port ${port}`);
}

One of the most critical engineering challenges in voice AI is handling human interruptions. If the AI is in the middle of speaking and the user says "Wait, how much does that cost?", the AI must shut up instantly.

To achieve this:

< 50ms. clear signal is immediately dispatched to Twilio to purge the pending audio playback buffer.

// Handling real-time human barge-in
function handleUserInterruption(ws: WebSocket, streamSid: string) {
  // Purge Twilio's audio queue immediately
  const clearPayload = JSON.stringify({
    event: 'clear',
    streamSid: streamSid
  });

  ws.send(clearPayload);

  // Abort pending TTS stream chunk generator
  ttsStreamController.abort();
}

In our Hikari Real Estate AI Solution, the voice agent qualifies buyer budget and timeline, books an executive site tour, and dispatches project brochures over WhatsApp in under 2 seconds:

// Define structured function calling schemas
const qualificationTools = [
  {
    type: "function",
    function: {
      name: "confirm_site_visit_and_send_whatsapp",
      description: "Confirms property site visit and immediately sends location & e-brochure over WhatsApp.",
      parameters: {
        type: "object",
        properties: {
          buyerName: { type: "string" },
          phoneNumber: { type: "string" },
          unitPreference: { type: "string", enum: ["2BHK", "3BHK", "Villa"] },
          budgetLakhs: { type: "number" },
          visitDateTime: { type: "string" }
        },
        required: ["buyerName", "phoneNumber", "unitPreference", "visitDateTime"]
      }
    }
  }
];

Across over 50,000 live conversations handled across our AI Phone Agent Deployments and Enterprise AI Operations:

To test live interactive ROI benchmarks and voice latency cost estimators, explore our Web & AI Cost Calculator or connect directly with our engineering team at Hikari Webworks.

── more in #ai-agents 4 stories · sorted by recency
── more on @hikari webworks 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-we-built-a-sub-5…] indexed:0 read:3min 2026-10-10 · —