Create an AI Podcast Generator with ElevenLabs A developer published a step-by-step guide for building an AI podcast generator that converts raw script text into a publishable MP3 using the ElevenLabs text-to-speech REST API. The walkthrough provides a Python function and an equivalent curl command that POST to the /v1/text-to-speech endpoint with a voice ID and settings such as stability 0.75 and similarity_boost 0.85, then stitch the returned audio chunks together with ffmpeg. The author claims the pipeline can produce a 30-minute episode in under a minute. Podcasts have exploded in popularity, but producing a high‑quality episode still takes a lot of time—recording, editing, and finding a consistent voice. What if you could turn a script into a polished episode with a single API call? In this guide we’ll build an AI Podcast Generator that takes raw text, feeds it to a text‑to‑speech TTS engine, and spits out an MP3 ready for publishing. The star of the show is ElevenLabs , a cutting‑edge voice AI platform that offers realistic voice cloning and a simple REST API. By the end of this article you’ll have a runnable Python script and a handy curl example that can generate a 30‑minute episode in under a minute. Why ElevenLabs? • Natural‑sounding voices that rival human narrators • Easy-to‑use API with per‑character pricing free tier for testing • Voice cloning lets you keep the same host voice across episodes Ready to give your podcast a voice? Let’s dive in. | What you need | Why it matters | |---|---| | Python 3.8+ or Node.js if you prefer | To call the ElevenLabs API and stitch audio files | | ffmpeg installed and in your PATH | For concatenating multiple audio chunks into a single MP3 | | An ElevenLabs API key sign up here https://try.elevenlabs.io/kr07zfuqn1bp | Grants access to the TTS service | | Basic knowledge of HTTP requests | Needed to interact with the REST endpoint | If you don’t have ffmpeg yet, on macOS you can run brew install ffmpeg , and on Ubuntu sudo apt-get install ffmpeg . ELEVENLABS API KEY . Create a new folder and install the required Python packages: mkdir ai-podcast cd ai-podcast python -m venv .venv source .venv/bin/activate Windows: .venv\Scripts\activate pip install requests tqdm We'll also add a small helper to download the generated audio chunks: python utils.py import os import requests from tqdm import tqdm def download file url: str, dest: str : resp = requests.get url, stream=True resp.raise for status total = int resp.headers.get 'content-length', 0 with open dest, 'wb' as f, tqdm desc=os.path.basename dest , total=total, unit='iB', unit scale=True, unit divisor=1024, as bar: for data in resp.iter content chunk size=1024 : size = f.write data bar.update size ElevenLabs expects a JSON payload with the text, the voice ID you want to use, and optional settings like stability and similarity boost. Below is a minimal Python function that sends a request and returns the URL of the generated MP3. python eleven.py import os import json import requests ELEVEN API KEY = os.getenv "ELEVENLABS API KEY" BASE URL = "https://api.elevenlabs.io/v1" def text to speech text: str, voice id: str = "EXAVITQu4vr4xnSDxMaL" - str: """ Sends text to ElevenLabs and returns a temporary URL to the generated audio. """ url = f"{BASE URL}/text-to-speech/{voice id}" headers = { "xi-api-key": ELEVEN API KEY, "Content-Type": "application/json", } payload = { "text": text, "model id": "eleven monolingual v1", "voice settings": { "stability": 0.75, "similarity boost": 0.85 } } response = requests.post url, headers=headers, json=payload response.raise for status The API returns the raw audio bytes; we’ll write them to a file. return response.content curl If you prefer a quick test from the command line, here’s the equivalent curl call: curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/EXAVITQu4vr4xnSDxMaL" \ -H "xi-api-key: $ELEVENLABS API KEY" \ -H "Content-Type: application/json" \ -d '{ "text": "Welcome to the AI Podcast Generator tutorial.", "model id": "eleven monolingual v1", "voice settings": { "stability": 0.75, "similarity boost": 0.85 } }' \ --output episode intro.mp3 Replace EXAVITQu4vr4xnSDxMaL with the voice ID you want ElevenLabs provides a few default voices; you can also upload a custom clone . Most podcasts are longer than a single API call can comfortably handle the API caps at ~5 k characters per request . The typical approach is to split the script into logical sections intro, interview, outro and generate each chunk separately. Below is a simple orchestrator that does exactly that: python generate episode.py import os import json import subprocess from pathlib import Path from eleven import text to speech ---------------------------------------------------------------------- 1️⃣ Load your script plain text, one paragraph per line ---------------------------------------------------------------------- SCRIPT PATH = Path "script.txt" segments = SCRIPT PATH.read text encoding="utf-8" .split "\n\n" double newline = segment ---------------------------------------------------------------------- 2️⃣ Generate audio for each segment ---------------------------------------------------------------------- audio files = for i, segment in enumerate segments, start=1 : print f"Generating segment {i}/{len segments } …" audio bytes = text to speech segment.strip out path = Path f"segment {i:03}.mp3" out path.write bytes audio bytes audio files.append str out path ---------------------------------------------------------------------- 3️⃣ Concatenate with ffmpeg ---------------------------------------------------------------------- list file = "concat list.txt" with open list file, "w" as f: for fp in audio files: f.write f"file '{fp}'\n" final mp3 = "episode full.mp3" subprocess.run "ffmpeg", "-y", "-f", "concat", "-safe", "0", "-i", list file, "-c", "copy", final mp3 , check=True print f"\n✅ Episode assembled: {final mp3}" script.txt is split on double newlines. Feel free to adjust the delimiter to match your writing style. text to speech . The function returns raw MP3 bytes, which we save to segment .mp3 . ffmpeg reads a tiny manifest concat list.txt and concatenates the files without re‑encoding, preserving the original quality. A podcast isn’t just a voice; you probably want intro music, background ambience, or a short ad slot. The same ffmpeg concat method can merge any number of audio tracks. Here’s a quick example that adds a 5‑second intro music clip: Prepare a manifest that interleaves music and voice cat concat list.txt <