Podcasts have exploded in popularity, but producing a high‑quality episode still takes a lot of time—recording, editing, and finding a consistent voice. What if you could turn a script into a polished episode with a single API call? In this guide we’ll build an AI Podcast Generator that takes raw text, feeds it to a text‑to‑speech (TTS) engine, and spits out an MP3 ready for publishing.
The star of the show is ElevenLabs, a cutting‑edge voice AI platform that offers realistic voice cloning and a simple REST API. By the end of this article you’ll have a runnable Python script (and a handy curl example) that can generate a 30‑minute episode in under a minute.
Why ElevenLabs?
• Natural‑sounding voices that rival human narrators
• Easy-to‑use API with per‑character pricing (free tier for testing)
• Voice cloning lets you keep the same host voice across episodes
Ready to give your podcast a voice? Let’s dive in.
| What you need | Why it matters |
|---|---|
| Python 3.8+ (or Node.js if you prefer) | To call the ElevenLabs API and stitch audio files |
ffmpeg installed and in yourPATH |
For concatenating multiple audio chunks into a single MP3 |
| An ElevenLabs API key (sign up here ) | Grants access to the TTS service |
| Basic knowledge of HTTP requests | Needed to interact with the REST endpoint |
If you don’t have ffmpeg yet, on macOS you can run brew install ffmpeg, and on Ubuntu sudo apt-get install ffmpeg.
ELEVENLABS_API_KEY.
Create a new folder and install the required Python packages:
mkdir ai-podcast
cd ai-podcast
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install requests tqdm
We'll also add a small helper to download the generated audio chunks:
import os
import requests
from tqdm import tqdm
def download_file(url: str, dest: str):
resp = requests.get(url, stream=True)
resp.raise_for_status()
total = int(resp.headers.get('content-length', 0))
with open(dest, 'wb') as f, tqdm(
desc=os.path.basename(dest),
total=total,
unit='iB',
unit_scale=True,
unit_divisor=1024,
) as bar:
for data in resp.iter_content(chunk_size=1024):
size = f.write(data)
bar.update(size)
ElevenLabs expects a JSON payload with the text, the voice ID you want to use, and optional settings like stability and similarity boost. Below is a minimal Python function that sends a request and returns the URL of the generated MP3.
import os
import json
import requests
ELEVEN_API_KEY = os.getenv("ELEVENLABS_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"
def text_to_speech(text: str, voice_id: str = "EXAVITQu4vr4xnSDxMaL") -> str:
"""
Sends `text` to ElevenLabs and returns a temporary URL to the generated audio.
"""
url = f"{BASE_URL}/text-to-speech/{voice_id}"
headers = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json",
}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1",
"voice_settings": {
"stability": 0.75,
"similarity_boost": 0.85
}
}
response = requests.post(url, headers=headers, json=payload)
response.raise_for_status()
return response.content
curl
If you prefer a quick test from the command line, here’s the equivalent curl call:
curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/EXAVITQu4vr4xnSDxMaL" \
-H "xi-api-key: $ELEVENLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"text": "Welcome to the AI Podcast Generator tutorial.",
"model_id": "eleven_monolingual_v1",
"voice_settings": { "stability": 0.75, "similarity_boost": 0.85 }
}' \
--output episode_intro.mp3
Replace EXAVITQu4vr4xnSDxMaL with the voice ID you want (ElevenLabs provides a few default voices; you can also upload a custom clone).
Most podcasts are longer than a single API call can comfortably handle (the API caps at ~5 k characters per request). The typical approach is to split the script into logical sections (intro, interview, outro) and generate each chunk separately. Below is a simple orchestrator that does exactly that:
import os
import json
import subprocess
from pathlib import Path
from eleven import text_to_speech
SCRIPT_PATH = Path("script.txt")
segments = SCRIPT_PATH.read_text(encoding="utf-8").split("\n\n") # double newline = segment
audio_files = []
for i, segment in enumerate(segments, start=1):
print(f"Generating segment {i}/{len(segments)} …")
audio_bytes = text_to_speech(segment.strip())
out_path = Path(f"segment_{i:03}.mp3")
out_path.write_bytes(audio_bytes)
audio_files.append(str(out_path))
list_file = "concat_list.txt"
with open(list_file, "w") as f:
for fp in audio_files:
f.write(f"file '{fp}'\n")
final_mp3 = "episode_full.mp3"
subprocess.run(
["ffmpeg", "-y", "-f", "concat", "-safe", "0", "-i", list_file,
"-c", "copy", final_mp3],
check=True
)
print(f"\n✅ Episode assembled: {final_mp3}")
script.txt) is split on double newlines. Feel free to adjust the delimiter to match your writing style.
text_to_speech. The function returns raw MP3 bytes, which we save to segment_###.mp3.
ffmpeg reads a tiny manifest (concat_list.txt) and concatenates the files without re‑encoding, preserving the original quality.
A podcast isn’t just a voice; you probably want intro music, background ambience, or a short ad slot. The same ffmpeg concat method can merge any number of audio tracks. Here’s a quick example that adds a 5‑second intro music clip:
cat > concat_list.txt <<EOF
file 'intro_music.mp3'
file 'segment_001.mp3'
file 'segment_002.mp3'
file 'outro_music.mp3'
EOF
ffmpeg -y -f concat -safe 0 -i concat_list.txt -c copy final_podcast.mp3
Make sure your music files have the same sample rate and channel layout as the TTS output (44.1 kHz, stereo) to avoid re‑encoding.
You can run the script locally, but for a production‑grade pipeline you’ll likely want a serverless function or a small Flask API. Below is a minimal Flask wrapper that accepts a JSON payload with a script field and returns a signed URL to the generated episode (using a temporary S3 bucket, for example). The core logic stays the same—just move the code from generate_episode.py into a function.
from flask import Flask, request, jsonify
from eleven import text_to_speech
import boto3, os, uuid, subprocess
app = Flask(__name__)
s3 = boto3.client("s3")
BUCKET = os.getenv("S3_BUCKET")
def build_episode(script: str) -> str:
segments = script.split("\n\n")
tmp_dir = f"/tmp/{uuid.uuid4()}"
os.makedirs(tmp_dir, exist_ok=True)
audio_paths = []
for i, seg in enumerate(segments, 1):
audio = text_to_speech(seg.strip())
path = f"{tmp_dir}/seg_{i:03}.mp3"
open(path, "wb").write(audio)
audio_paths.append(path)
list_file = f"{tmp_dir}/list.txt"
with open(list_file, "w") as f:
for p in audio_paths:
f.write(f"file '{p}'\n")
final_path = f"{tmp_dir}/episode.mp3"
subprocess.run(
["ffmpeg", "-y", "-f", "concat", "-safe", "0", "-i", list_file,
"-c", "copy", final_path],
check=True
)
key = f"episodes/{uuid.uuid4()}.mp3"
s3.upload_file(final_path, BUCKET, key, ExtraArgs={"ACL": "public-read"})
return f"https://{BUCKET}.s3.amazonaws.com/{key}"
@app.route("/generate", methods=["POST"])
def generate():
data = request.get_json()
if not data or "script" not in data:
return jsonify({"error": "Missing script"}), 400
url = build_episode(data["script"])
return jsonify({"episode_url": url})
if __name__ == "__main__":
app.run(debug=True)
Deploy this to a platform like Render, Fly.io, or AWS Lambda + API Gateway and you have a fully automated podcast generator that anyone can call via a simple HTTP request.
| Tip | Why it matters |
|---|---|
| Keep sentences under 150 characters | ElevenLabs handles short bursts more naturally; longer sentences can produce slight breath artifacts. |
Add s with , or ... |
The engine interprets punctuation as breath or cues, giving a more human rhythm. |
| Use the same voice ID for every episode | Consistency builds brand identity. Clone your own voice if you want a unique host. |
| Test stability & similarity settings | Higher stability yields smoother speech; similarity boost makes the voice sound more like the reference. |
Feel free to experiment with the voice_settings payload – the API docs (linked from the ElevenLabs dashboard) provide a nice interactive playground.
You now have a complete end‑to‑end workflow:
All of this is powered by ElevenLabs, whose realistic voice cloning makes the final product sound like a professional narrator rather than a robotic read‑out.
Give it a spin, tweak the voice settings, and start churning out episodes without ever stepping into a recording booth.
Ready to give your podcast a voice? Sign up at https://try.elevenlabs.io/kr07zfuqn1bp and start generating today!