If you’ve ever tried to build a voice assistant, you know the biggest headache is getting natural‑sounding speech. Traditional TTS engines either sound robotic or require massive amounts of data and compute. ElevenLabs solves both problems with a cloud API that delivers studio‑grade speech and even lets you clone a voice in minutes. The best part? You can start using it from a simple Node.js script and scale up to full‑blown conversational agents.
Quick tip: Sign up through this affiliate link – it gives you a free credit to experiment: https://try.elevenlabs.io/kr07zfuqn1bp
At a high level, our AI voice assistant will consist of three parts:
All the heavy lifting is done by ElevenLabs, so you can focus on the conversational logic.
mkdir ai-voice-assistant && cd ai-voice-assistant
npm init -y
npm install express axios cors dotenv
Create a .env file to keep your API key safe:
ELEVENLABS_API_KEY=your_elevenlabs_api_key_here
PORT=3000
Note: Grab your API key from the ElevenLabs dashboard after signing up via https://try.elevenlabs.io/kr07zfuqn1bp.
Below is a minimal server that receives a JSON payload like { "text": "Hello, world!" } and returns a base64‑encoded audio string.
// server.js
require('dotenv').config();
const express = require('express');
const axios = require('axios');
const cors = require('cors');
const app = express();
app.use(cors());
app.use(express.json());
const ELEVENLABS_API_KEY = process.env.ELEVENLABS_API_KEY;
const VOICE_ID = 'EXAVITQu4vr4xnSDxMaL'; // default voice; replace with your cloned voice ID
app.post('/synthesize', async (req, res) => {
const { text } = req.body;
if (!text) return res.status(400).json({ error: 'Missing text' });
try {
const response = await axios({
method: 'post',
url: `https://api.elevenlabs.io/v1/text-to-speech/${VOICE_ID}`,
headers: {
'xi-api-key': ELEVENLABS_API_KEY,
'Content-Type': 'application/json',
'Accept': 'audio/mpeg',
},
data: {
text,
voice_settings: {
stability: 0.75,
similarity_boost: 0.85,
},
},
responseType: 'arraybuffer',
});
const base64Audio = Buffer.from(response.data, 'binary').toString('base64');
res.json({ audio: base64Audio });
} catch (err) {
console.error('ElevenLabs error:', err.response?.data || err.message);
res.status(500).json({ error: 'TTS failed' });
}
});
app.listen(process.env.PORT, () => {
console.log(`🚀 Server listening on http://localhost:${process.env.PORT}`);
});
audio/mpeg (MP3) because it streams easily in browsers.
Run the server:
node server.js
Create a simple HTML page that uses the Web Speech API for STT and fetches the synthesized audio from our server.
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<title>AI Voice Assistant Demo</title>
<style>
body { font-family: sans-serif; margin: 2rem; }
button { padding: .5rem 1rem; margin-top: 1rem; }
</style>
</head>
<body>
<h1>Talk to the AI</h1>
<button id="talkBtn">Start Listening</button>
<p id="transcript"></p>
<audio id="replyAudio" controls></audio>
<script>
const talkBtn = document.getElementById('talkBtn');
const transcriptEl = document.getElementById('transcript');
const audioEl = document.getElementById('replyAudio');
const SpeechRecognition = window.SpeechRecognition || window.webkitSpeechRecognition;
const recognizer = new SpeechRecognition();
recognizer.lang = 'en-US';
recognizer.interimResults = false;
talkBtn.onclick = () => {
recognizer.start();
talkBtn.disabled = true;
talkBtn.textContent = 'Listening...';
};
recognizer.onresult = async (event) => {
const spoken = event.results[0][0].transcript;
transcriptEl.textContent = `You said: "${spoken}"`;
talkBtn.disabled = false;
talkBtn.textContent = 'Start Listening';
// TODO: replace this with your LLM call; for demo we just echo back
const reply = `You just said: ${spoken}`;
const resp = await fetch('http://localhost:3000/synthesize', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ text: reply })
});
const data = await resp.json();
audioEl.src = `data:audio/mpeg;base64,${data.audio}`;
audioEl.play();
};
recognizer.onerror = (e) => {
console.error(e);
talkBtn.disabled = false;
talkBtn.textContent = 'Start Listening';
};
</script>
</body>
</html>
Open index.html in a browser, click Start Listening, and speak. The assistant will repeat what you said using ElevenLabs‑generated speech.
One of ElevenLabs’ standout features is the ability to create a custom voice from as little as 30 seconds of audio. Here’s a quick rundown:
VOICE_ID constant in server.js with this new ID.
Now your assistant will speak with your voice, your colleague’s voice, or even a fictional character’s voice—all without the need for a deep learning pipeline.
The demo above simply echoes the user’s input. In a production app you’ll want a language model to generate meaningful replies. Here’s a minimal example using OpenAI’s gpt-3.5-turbo:
// add to server.js (install openai: npm i openai)
const { OpenAI } = require('openai');
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
async function getChatResponse(message) {
const completion = await openai.chat.completions.create({
model: 'gpt-3.5-turbo',
messages: [{ role: 'user', content: message }],
});
return completion.choices[0].message.content.trim();
}
// Inside /synthesize route, replace the static reply:
const reply = await getChatResponse(text);
Now the flow becomes:
| Issue | Likely Cause | Fix |
|---|---|---|
| 401 Unauthorized from ElevenLabs | Wrong or missing API key | Verify ELEVENLABS_API_KEY in.env and that the key is active. |
| No audio returned, empty base64 string | Accept: audio/mpeg missing or responseType not set |
Ensure responseType: 'arraybuffer' andAccept header are present. |
| Voice sounds robotic | Using default voice with low stability | Increase stability (0.7‑0.9) andsimilarity_boost . |
| Long latency (>3 s) | Large text payload or network throttling | Split long paragraphs into smaller chunks and synthesize sequentially. |
When you’re ready to go public, you can push the Express app to services like Vercel, Render, or Railway. Because the TTS endpoint streams binary data, make sure the platform supports responseType: 'arraybuffer'. You’ll also need to set environment variables ( ELEVENLABS_API_KEY, OPENAI_API_KEY, etc.) in the dashboard of your chosen host.
Building an AI voice assistant is now a matter of stitching together a few well‑documented APIs:
The heavy lifting—high‑quality synthesis and voice cloning—is handled by ElevenLabs, letting you ship a polished product in days instead of weeks.
Ready to give your assistant a human voice? Try ElevenLabs today via this link and get a free credit to start experimenting: https://try.elevenlabs.io/kr07zfuqn1bp