Create an AI Voice Assistant with ElevenLabs and Node.js A developer published a tutorial showing how to build an AI voice assistant using the ElevenLabs text-to-speech API and Node.js, with a minimal Express server that accepts text and returns base64-encoded MP3 audio. The walkthrough pairs ElevenLabs synthesis with the browser's Web Speech API for speech recognition and includes voice settings for stability and similarity boost. If you’ve ever tried to build a voice assistant, you know the biggest headache is getting natural‑sounding speech. Traditional TTS engines either sound robotic or require massive amounts of data and compute. ElevenLabs solves both problems with a cloud API that delivers studio‑grade speech and even lets you clone a voice in minutes. The best part? You can start using it from a simple Node.js script and scale up to full‑blown conversational agents. Quick tip: Sign up through this affiliate link – it gives you a free credit to experiment: https://try.elevenlabs.io/kr07zfuqn1bp https://try.elevenlabs.io/kr07zfuqn1bp At a high level, our AI voice assistant will consist of three parts: All the heavy lifting is done by ElevenLabs, so you can focus on the conversational logic. Create a fresh folder mkdir ai-voice-assistant && cd ai-voice-assistant Initialize a Node project npm init -y Install dependencies npm install express axios cors dotenv Create a .env file to keep your API key safe: ELEVENLABS API KEY=your elevenlabs api key here PORT=3000 Note: Grab your API key from the ElevenLabs dashboard after signing up via https://try.elevenlabs.io/kr07zfuqn1bp https://try.elevenlabs.io/kr07zfuqn1bp . Below is a minimal server that receives a JSON payload like { "text": "Hello, world " } and returns a base64‑encoded audio string. js // server.js require 'dotenv' .config ; const express = require 'express' ; const axios = require 'axios' ; const cors = require 'cors' ; const app = express ; app.use cors ; app.use express.json ; const ELEVENLABS API KEY = process.env.ELEVENLABS API KEY; const VOICE ID = 'EXAVITQu4vr4xnSDxMaL'; // default voice; replace with your cloned voice ID app.post '/synthesize', async req, res = { const { text } = req.body; if text return res.status 400 .json { error: 'Missing text' } ; try { const response = await axios { method: 'post', url: https://api.elevenlabs.io/v1/text-to-speech/${VOICE ID} , headers: { 'xi-api-key': ELEVENLABS API KEY, 'Content-Type': 'application/json', 'Accept': 'audio/mpeg', }, data: { text, voice settings: { stability: 0.75, similarity boost: 0.85, }, }, responseType: 'arraybuffer', } ; const base64Audio = Buffer.from response.data, 'binary' .toString 'base64' ; res.json { audio: base64Audio } ; } catch err { console.error 'ElevenLabs error:', err.response?.data || err.message ; res.status 500 .json { error: 'TTS failed' } ; } } ; app.listen process.env.PORT, = { console.log 🚀 Server listening on http://localhost:${process.env.PORT} ; } ; audio/mpeg MP3 because it streams easily in browsers. Run the server: node server.js Create a simple HTML page that uses the Web Speech API for STT and fetches the synthesized audio from our server. < DOCTYPE html