Integrating ElevenLabs with Next.js: Step-by-Step Guide A developer published a step-by-step guide for integrating ElevenLabs' text-to-speech API into a Next.js application, using a server-side API route to keep the ElevenLabs API key hidden from the browser. The walkthrough covers creating a Next.js app, storing credentials in environment variables, proxying synthesis requests through a Next.js API route that returns raw MP3 audio, and playing the streamed result in a React component. If you’ve ever wanted to add realistic, AI‑generated voice to a web app—whether it’s a podcast generator, an interactive story, or a voice‑enabled chatbot—ElevenLabs is one of the most powerful text‑to‑speech TTS engines on the market today. Its neural models produce natural‑sounding speech, support multiple languages, and even let you clone a voice with a few minutes of audio. Next.js, with its hybrid rendering capabilities and built‑in API routes, makes it incredibly easy to call external services like ElevenLabs from both the client and the server. In this guide we’ll walk through a full‑stack implementation: a simple UI where users type text, hit “Speak”, and hear the audio streamed back instantly. npx create-next-app@latest my-voice-app Create a .env.local file at the root of your project and add the following: ELEVENLABS API KEY=your-elevenlabs-api-key ELEVENLABS VOICE ID=your-default-voice-id e.g., “EXAVITQu4vr4xnSDxMaL” Tip: Keep your API key secret—Next.js automatically exposes only variables prefixed with NEXT PUBLIC to the browser. Since we’ll call ElevenLabs from the server, we don’t need to expose it. Next.js API routes run on the server, perfect for keeping the API key hidden. Create a new file: pages/api/speak.ts . python import type { NextApiRequest, NextApiResponse } from 'next'; import fetch from 'node-fetch'; export default async function handler req: NextApiRequest, res: NextApiResponse { if req.method == 'POST' { return res.status 405 .json { error: 'Method not allowed' } ; } const { text, voiceId } = req.body; if text { return res.status 400 .json { error: 'Missing text payload' } ; } const voice = voiceId || process.env.ELEVENLABS VOICE ID; const apiKey = process.env.ELEVENLABS API KEY; try { const response = await fetch https://api.elevenlabs.io/v1/text-to-speech/${voice} , { method: 'POST', headers: { 'Content-Type': 'application/json', 'xi-api-key': apiKey , }, body: JSON.stringify { text, voice settings: { stability: 0.5, similarity boost: 0.75, }, } , } ; if response.ok { const err = await response.text ; throw new Error ElevenLabs error: ${err} ; } // The API returns raw audio mp3 bytes const audioBuffer = await response.arrayBuffer ; // Stream the audio back to the client res.setHeader 'Content-Type', 'audio/mpeg' ; res.send Buffer.from audioBuffer ; } catch error: any { console.error error ; res.status 500 .json { error: error.message } ; } } What’s happening? POST request with text the script and an optional voiceId . Create a component at components/VoiceSynthesizer.tsx : js import { useState, useRef } from 'react'; export default function VoiceSynthesizer { const text, setText = useState '' ; const loading, setLoading = useState false ; const audioRef = useRef