{"slug": "create-an-ai-voice-assistant-with-elevenlabs-and-node-js", "title": "Create an AI Voice Assistant with ElevenLabs and Node.js", "summary": "A developer published a tutorial showing how to build an AI voice assistant using the ElevenLabs text-to-speech API and Node.js, with a minimal Express server that accepts text and returns base64-encoded MP3 audio. The walkthrough pairs ElevenLabs synthesis with the browser's Web Speech API for speech recognition and includes voice settings for stability and similarity boost.", "body_md": "If you’ve ever tried to build a voice assistant, you know the biggest headache is getting natural‑sounding speech. Traditional TTS engines either sound robotic or require massive amounts of data and compute. ElevenLabs solves both problems with a cloud API that delivers studio‑grade speech and even lets you clone a voice in minutes. The best part? You can start using it from a simple Node.js script and scale up to full‑blown conversational agents.\n\n**Quick tip:** Sign up through this affiliate link – it gives you a free credit to experiment: [https://try.elevenlabs.io/kr07zfuqn1bp](https://try.elevenlabs.io/kr07zfuqn1bp)\n\nAt a high level, our AI voice assistant will consist of three parts:\n\nAll the heavy lifting is done by ElevenLabs, so you can focus on the conversational logic.\n\n```\n# Create a fresh folder\nmkdir ai-voice-assistant && cd ai-voice-assistant\n\n# Initialize a Node project\nnpm init -y\n\n# Install dependencies\nnpm install express axios cors dotenv\n```\n\nCreate a `.env` file to keep your API key safe:\n\n```\nELEVENLABS_API_KEY=your_elevenlabs_api_key_here\nPORT=3000\n```\n\n**Note:** Grab your API key from the ElevenLabs dashboard after signing up via [https://try.elevenlabs.io/kr07zfuqn1bp](https://try.elevenlabs.io/kr07zfuqn1bp).\n\nBelow is a minimal server that receives a JSON payload like `{ \"text\": \"Hello, world!\" }` and returns a base64‑encoded audio string.\n\n``` js\n// server.js\nrequire('dotenv').config();\nconst express = require('express');\nconst axios = require('axios');\nconst cors = require('cors');\n\nconst app = express();\napp.use(cors());\napp.use(express.json());\n\nconst ELEVENLABS_API_KEY = process.env.ELEVENLABS_API_KEY;\nconst VOICE_ID = 'EXAVITQu4vr4xnSDxMaL'; // default voice; replace with your cloned voice ID\n\napp.post('/synthesize', async (req, res) => {\n  const { text } = req.body;\n  if (!text) return res.status(400).json({ error: 'Missing text' });\n\n  try {\n    const response = await axios({\n      method: 'post',\n      url: `https://api.elevenlabs.io/v1/text-to-speech/${VOICE_ID}`,\n      headers: {\n        'xi-api-key': ELEVENLABS_API_KEY,\n        'Content-Type': 'application/json',\n        'Accept': 'audio/mpeg',\n      },\n      data: {\n        text,\n        voice_settings: {\n          stability: 0.75,\n          similarity_boost: 0.85,\n        },\n      },\n      responseType: 'arraybuffer',\n    });\n\n    const base64Audio = Buffer.from(response.data, 'binary').toString('base64');\n    res.json({ audio: base64Audio });\n  } catch (err) {\n    console.error('ElevenLabs error:', err.response?.data || err.message);\n    res.status(500).json({ error: 'TTS failed' });\n  }\n});\n\napp.listen(process.env.PORT, () => {\n  console.log(`🚀 Server listening on http://localhost:${process.env.PORT}`);\n});\n```\n\n`audio/mpeg` (MP3) because it streams easily in browsers.\nRun the server:\n\n```\nnode server.js\n```\n\nCreate a simple HTML page that uses the Web Speech API for STT and fetches the synthesized audio from our server.\n\n```\n<!DOCTYPE html>\n<html lang=\"en\">\n<head>\n  <meta charset=\"UTF-8\">\n  <title>AI Voice Assistant Demo</title>\n  <style>\n    body { font-family: sans-serif; margin: 2rem; }\n    button { padding: .5rem 1rem; margin-top: 1rem; }\n  </style>\n</head>\n<body>\n  <h1>Talk to the AI</h1>\n  <button id=\"talkBtn\">Start Listening</button>\n  <p id=\"transcript\"></p>\n  <audio id=\"replyAudio\" controls></audio>\n\n  <script>\n    const talkBtn = document.getElementById('talkBtn');\n    const transcriptEl = document.getElementById('transcript');\n    const audioEl = document.getElementById('replyAudio');\n\n    const SpeechRecognition = window.SpeechRecognition || window.webkitSpeechRecognition;\n    const recognizer = new SpeechRecognition();\n    recognizer.lang = 'en-US';\n    recognizer.interimResults = false;\n\n    talkBtn.onclick = () => {\n      recognizer.start();\n      talkBtn.disabled = true;\n      talkBtn.textContent = 'Listening...';\n    };\n\n    recognizer.onresult = async (event) => {\n      const spoken = event.results[0][0].transcript;\n      transcriptEl.textContent = `You said: \"${spoken}\"`;\n      talkBtn.disabled = false;\n      talkBtn.textContent = 'Start Listening';\n\n      // TODO: replace this with your LLM call; for demo we just echo back\n      const reply = `You just said: ${spoken}`;\n\n      const resp = await fetch('http://localhost:3000/synthesize', {\n        method: 'POST',\n        headers: { 'Content-Type': 'application/json' },\n        body: JSON.stringify({ text: reply })\n      });\n      const data = await resp.json();\n      audioEl.src = `data:audio/mpeg;base64,${data.audio}`;\n      audioEl.play();\n    };\n\n    recognizer.onerror = (e) => {\n      console.error(e);\n      talkBtn.disabled = false;\n      talkBtn.textContent = 'Start Listening';\n    };\n  </script>\n</body>\n</html>\n```\n\nOpen `index.html` in a browser, click **Start Listening**, and speak. The assistant will repeat what you said using ElevenLabs‑generated speech.\n\nOne of ElevenLabs’ standout features is the ability to create a custom voice from as little as 30 seconds of audio. Here’s a quick rundown:\n\n`VOICE_ID` constant in `server.js` with this new ID.\nNow your assistant will speak with *your* voice, your colleague’s voice, or even a fictional character’s voice—all without the need for a deep learning pipeline.\n\nThe demo above simply echoes the user’s input. In a production app you’ll want a language model to generate meaningful replies. Here’s a minimal example using OpenAI’s `gpt-3.5-turbo`:\n\n``` js\n// add to server.js (install openai: npm i openai)\nconst { OpenAI } = require('openai');\nconst openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });\n\nasync function getChatResponse(message) {\n  const completion = await openai.chat.completions.create({\n    model: 'gpt-3.5-turbo',\n    messages: [{ role: 'user', content: message }],\n  });\n  return completion.choices[0].message.content.trim();\n}\n\n// Inside /synthesize route, replace the static reply:\nconst reply = await getChatResponse(text);\n```\n\nNow the flow becomes:\n\n| Issue | Likely Cause | Fix | \n|---|---|---|\n| 401 Unauthorized from ElevenLabs | Wrong or missing API key | Verify `ELEVENLABS_API_KEY` in`.env` and that the key is active. | \n| No audio returned, empty base64 string | `Accept: audio/mpeg` missing or responseType not set | Ensure `responseType: 'arraybuffer'` and`Accept` header are present. | \n| Voice sounds robotic | Using default voice with low stability | Increase `stability` (0.7‑0.9) and`similarity_boost` . | \n| Long latency (>3 s) | Large text payload or network throttling | Split long paragraphs into smaller chunks and synthesize sequentially. | \n\nWhen you’re ready to go public, you can push the Express app to services like Vercel, Render, or Railway. Because the TTS endpoint streams binary data, make sure the platform supports `responseType: 'arraybuffer'`. You’ll also need to set environment variables (` ELEVENLABS_API_KEY`, `OPENAI_API_KEY`, etc.) in the dashboard of your chosen host.\n\nBuilding an AI voice assistant is now a matter of stitching together a few well‑documented APIs:\n\nThe heavy lifting—high‑quality synthesis and voice cloning—is handled by ElevenLabs, letting you ship a polished product in days instead of weeks.\n\n**Ready to give your assistant a human voice?** Try ElevenLabs today via this link and get a free credit to start experimenting: [https://try.elevenlabs.io/kr07zfuqn1bp](https://try.elevenlabs.io/kr07zfuqn1bp)", "url": "https://wpnews.pro/news/create-an-ai-voice-assistant-with-elevenlabs-and-node-js", "canonical_source": "https://dev.to/voice_developer/create-an-ai-voice-assistant-with-elevenlabs-and-nodejs-2nci", "published_at": "2026-10-02 22:49:28+00:00", "updated_at": "2026-10-02 23:08:08.880466+00:00", "lang": "en", "topics": ["ai-tools", "natural-language-processing", "ai-products"], "entities": ["ElevenLabs", "Node.js", "Express", "Web Speech API"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/create-an-ai-voice-assistant-with-elevenlabs-and-node-js", "markdown": "https://wpnews.pro/news/create-an-ai-voice-assistant-with-elevenlabs-and-node-js.md", "text": "https://wpnews.pro/news/create-an-ai-voice-assistant-with-elevenlabs-and-node-js.txt", "jsonld": "https://wpnews.pro/news/create-an-ai-voice-assistant-with-elevenlabs-and-node-js.jsonld"}}