When building conversational AI phone agents for real-world production environments (such as after-hours dental clinic receptionists or real estate lead qualifiers), latency is everything.
In human conversation, the natural between speakers averages 200ms to 400ms. If an AI telephone system takes 1.5 to 3 seconds to respond:
In this tutorial, we share the exact streaming architecture we engineered at Hikari Webworks to achieve sub-500ms end-to-end voice latency in production.
Sequential HTTP request-response cycles (Record audio -> Send POST -> Wait for LLM -> Generate MP3 -> Play audio) add at least 2.5 seconds of artificial delay.
To achieve sub-500ms, the entire pipeline must operate over full-duplex bi-directional WebSockets:
[Phone Caller]
│ (8kHz raw audio stream over PSTN/SIP)
▼
[Twilio Media Stream WebSocket]
│
├── (80ms) ──► Deepgram Nova-2 (Live Streaming STT)
│
├── (140ms) ─► Streaming LLM (Groq Llama 3.3 70B / Claude 3.5 Haiku)
│
├── (90ms) ──► Ultra-Fast TTS (Cartesia Sonic / ElevenLabs Turbo)
│
└── (40ms) ──► Raw Mu-Law Audio Chunk Streaming back to Caller
Total Turnaround Time: ~350ms - 480ms (Sub-500ms Real-Time Loop)
When a call connects, Twilio opens a WebSocket connection and streams 8kHz 16-bit Mu-law audio packets encoded in base64:
// server/voice-agent.ts
import { WebSocketServer, WebSocket } from 'ws';
interface TwilioMediaMessage {
event: 'connected' | 'start' | 'media' | 'stop' | 'mark' | 'clear';
streamSid?: string;
media?: {
payload: string; // Base64 encoded 8kHz mu-law audio
timestamp: string;
chunk: string;
};
}
export function initializeVoiceServer(port: number = 8080) {
const wss = new WebSocketServer({ port });
wss.on('connection', (ws: WebSocket) => {
let streamSid = '';
ws.on('message', async (data: string) => {
const msg: TwilioMediaMessage = JSON.parse(data);
if (msg.event === 'start' && msg.streamSid) {
streamSid = msg.streamSid;
console.log(`[Twilio] Call stream active: ${streamSid}`);
}
if (msg.event === 'media' && msg.media) {
// Decode raw audio frame and stream immediately to Speech-to-Text
const rawChunk = Buffer.from(msg.media.payload, 'base64');
liveSttStream.send(rawChunk);
}
if (msg.event === 'stop') {
console.log(`[Twilio] Call ended: ${streamSid}`);
}
});
});
console.log(`AI Voice Gateway listening on port ${port}`);
}
One of the most critical engineering challenges in voice AI is handling human interruptions. If the AI is in the middle of speaking and the user says "Wait, how much does that cost?", the AI must shut up instantly.
To achieve this:
< 50ms. clear signal is immediately dispatched to Twilio to purge the pending audio playback buffer.
// Handling real-time human barge-in
function handleUserInterruption(ws: WebSocket, streamSid: string) {
// Purge Twilio's audio queue immediately
const clearPayload = JSON.stringify({
event: 'clear',
streamSid: streamSid
});
ws.send(clearPayload);
// Abort pending TTS stream chunk generator
ttsStreamController.abort();
}
In our Hikari Real Estate AI Solution, the voice agent qualifies buyer budget and timeline, books an executive site tour, and dispatches project brochures over WhatsApp in under 2 seconds:
// Define structured function calling schemas
const qualificationTools = [
{
type: "function",
function: {
name: "confirm_site_visit_and_send_whatsapp",
description: "Confirms property site visit and immediately sends location & e-brochure over WhatsApp.",
parameters: {
type: "object",
properties: {
buyerName: { type: "string" },
phoneNumber: { type: "string" },
unitPreference: { type: "string", enum: ["2BHK", "3BHK", "Villa"] },
budgetLakhs: { type: "number" },
visitDateTime: { type: "string" }
},
required: ["buyerName", "phoneNumber", "unitPreference", "visitDateTime"]
}
}
}
];
Across over 50,000 live conversations handled across our AI Phone Agent Deployments and Enterprise AI Operations:
To test live interactive ROI benchmarks and voice latency cost estimators, explore our Web & AI Cost Calculator or connect directly with our engineering team at Hikari Webworks.