# How We Built a Sub-500ms Real-Time AI Voice Agent with TypeScript, Twilio & WebSockets

> Source: <https://dev.to/devtrivedi/how-we-built-a-sub-500ms-real-time-ai-voice-agent-with-typescript-twilio-websockets-3hlg>
> Published: 2026-10-10 05:30:00+00:00

When building conversational AI phone agents for real-world production environments (such as after-hours dental clinic receptionists or real estate lead qualifiers), **latency is everything**.

In human conversation, the natural pause between speakers averages **200ms to 400ms**. If an AI telephone system takes 1.5 to 3 seconds to respond:

In this tutorial, we share the exact streaming architecture we engineered at [Hikari Webworks](https://hikariwebworks.studio) to achieve **sub-500ms end-to-end voice latency** in production.

Sequential HTTP request-response cycles (`Record audio -> Send POST -> Wait for LLM -> Generate MP3 -> Play audio`) add at least 2.5 seconds of artificial delay.

To achieve sub-500ms, the entire pipeline must operate over **full-duplex bi-directional WebSockets**:

```
[Phone Caller] 
      │ (8kHz raw audio stream over PSTN/SIP)
      ▼
[Twilio Media Stream WebSocket]
      │
      ├── (80ms) ──► Deepgram Nova-2 (Live Streaming STT)
      │
      ├── (140ms) ─► Streaming LLM (Groq Llama 3.3 70B / Claude 3.5 Haiku)
      │
      ├── (90ms) ──► Ultra-Fast TTS (Cartesia Sonic / ElevenLabs Turbo)
      │
      └── (40ms) ──► Raw Mu-Law Audio Chunk Streaming back to Caller

Total Turnaround Time: ~350ms - 480ms (Sub-500ms Real-Time Loop)
```

When a call connects, Twilio opens a WebSocket connection and streams 8kHz 16-bit Mu-law audio packets encoded in base64:

``` js
// server/voice-agent.ts
import { WebSocketServer, WebSocket } from 'ws';

interface TwilioMediaMessage {
  event: 'connected' | 'start' | 'media' | 'stop' | 'mark' | 'clear';
  streamSid?: string;
  media?: {
    payload: string; // Base64 encoded 8kHz mu-law audio
    timestamp: string;
    chunk: string;
  };
}

export function initializeVoiceServer(port: number = 8080) {
  const wss = new WebSocketServer({ port });

  wss.on('connection', (ws: WebSocket) => {
    let streamSid = '';

    ws.on('message', async (data: string) => {
      const msg: TwilioMediaMessage = JSON.parse(data);

      if (msg.event === 'start' && msg.streamSid) {
        streamSid = msg.streamSid;
        console.log(`[Twilio] Call stream active: ${streamSid}`);
      }

      if (msg.event === 'media' && msg.media) {
        // Decode raw audio frame and stream immediately to Speech-to-Text
        const rawChunk = Buffer.from(msg.media.payload, 'base64');
        liveSttStream.send(rawChunk);
      }

      if (msg.event === 'stop') {
        console.log(`[Twilio] Call ended: ${streamSid}`);
      }
    });
  });

  console.log(`AI Voice Gateway listening on port ${port}`);
}
```

One of the most critical engineering challenges in voice AI is handling human interruptions. If the AI is in the middle of speaking and the user says *"Wait, how much does that cost?"*, the AI must shut up instantly.

To achieve this:

`< 50ms`.` clear` signal is immediately dispatched to Twilio to purge the pending audio playback buffer.

```
// Handling real-time human barge-in
function handleUserInterruption(ws: WebSocket, streamSid: string) {
  // Purge Twilio's audio queue immediately
  const clearPayload = JSON.stringify({
    event: 'clear',
    streamSid: streamSid
  });

  ws.send(clearPayload);

  // Abort pending TTS stream chunk generator
  ttsStreamController.abort();
}
```

In our [Hikari Real Estate AI Solution](https://hikariwebworks.studio/solutions/real-estate), the voice agent qualifies buyer budget and timeline, books an executive site tour, and dispatches project brochures over WhatsApp in under 2 seconds:

``` js
// Define structured function calling schemas
const qualificationTools = [
  {
    type: "function",
    function: {
      name: "confirm_site_visit_and_send_whatsapp",
      description: "Confirms property site visit and immediately sends location & e-brochure over WhatsApp.",
      parameters: {
        type: "object",
        properties: {
          buyerName: { type: "string" },
          phoneNumber: { type: "string" },
          unitPreference: { type: "string", enum: ["2BHK", "3BHK", "Villa"] },
          budgetLakhs: { type: "number" },
          visitDateTime: { type: "string" }
        },
        required: ["buyerName", "phoneNumber", "unitPreference", "visitDateTime"]
      }
    }
  }
];
```

Across over 50,000 live conversations handled across our [AI Phone Agent Deployments](https://hikariwebworks.studio/services/phone-agents) and [Enterprise AI Operations](https://hikariwebworks.studio/solutions/ai-operations):

To test live interactive ROI benchmarks and voice latency cost estimators, explore our [Web & AI Cost Calculator](https://hikariwebworks.studio/cost-calculator) or connect directly with our engineering team at [Hikari Webworks](https://hikariwebworks.studio).
