{"slug": "create-an-ai-podcast-generator-with-elevenlabs", "title": "Create an AI Podcast Generator with ElevenLabs", "summary": "A developer published a step-by-step guide for building an AI podcast generator that converts raw script text into a publishable MP3 using the ElevenLabs text-to-speech REST API. The walkthrough provides a Python function and an equivalent curl command that POST to the /v1/text-to-speech endpoint with a voice ID and settings such as stability 0.75 and similarity_boost 0.85, then stitch the returned audio chunks together with ffmpeg. The author claims the pipeline can produce a 30-minute episode in under a minute.", "body_md": "Podcasts have exploded in popularity, but producing a high‑quality episode still takes a lot of time—recording, editing, and finding a consistent voice. What if you could turn a script into a polished episode with a single API call? In this guide we’ll build an **AI Podcast Generator** that takes raw text, feeds it to a text‑to‑speech (TTS) engine, and spits out an MP3 ready for publishing.  \n\nThe star of the show is **ElevenLabs**, a cutting‑edge voice AI platform that offers realistic voice cloning and a simple REST API. By the end of this article you’ll have a runnable Python script (and a handy `curl` example) that can generate a 30‑minute episode in under a minute.\n\n**Why ElevenLabs?**\n\n• Natural‑sounding voices that rival human narrators\n\n• Easy-to‑use API with per‑character pricing (free tier for testing)\n\n• Voice cloning lets you keep the same host voice across episodes  \n\nReady to give your podcast a voice? Let’s dive in.\n\n| What you need | Why it matters | \n|---|---|\n| Python 3.8+ (or Node.js if you prefer) | To call the ElevenLabs API and stitch audio files | \n| `ffmpeg` installed and in your`PATH` | For concatenating multiple audio chunks into a single MP3 | \n| An ElevenLabs API key (sign up [here](https://try.elevenlabs.io/kr07zfuqn1bp) ) | Grants access to the TTS service | \n| Basic knowledge of HTTP requests | Needed to interact with the REST endpoint | \n\nIf you don’t have `ffmpeg` yet, on macOS you can run `brew install ffmpeg`, and on Ubuntu `sudo apt-get install ffmpeg`.\n\n`ELEVENLABS_API_KEY`.\nCreate a new folder and install the required Python packages:\n\n```\nmkdir ai-podcast\ncd ai-podcast\npython -m venv .venv\nsource .venv/bin/activate   # Windows: .venv\\Scripts\\activate\npip install requests tqdm\n```\n\nWe'll also add a small helper to download the generated audio chunks:\n\n``` python\n# utils.py\nimport os\nimport requests\nfrom tqdm import tqdm\n\ndef download_file(url: str, dest: str):\n    resp = requests.get(url, stream=True)\n    resp.raise_for_status()\n    total = int(resp.headers.get('content-length', 0))\n    with open(dest, 'wb') as f, tqdm(\n        desc=os.path.basename(dest),\n        total=total,\n        unit='iB',\n        unit_scale=True,\n        unit_divisor=1024,\n    ) as bar:\n        for data in resp.iter_content(chunk_size=1024):\n            size = f.write(data)\n            bar.update(size)\n```\n\nElevenLabs expects a JSON payload with the text, the voice ID you want to use, and optional settings like stability and similarity boost. Below is a minimal Python function that sends a request and returns the URL of the generated MP3.\n\n``` python\n# eleven.py\nimport os\nimport json\nimport requests\n\nELEVEN_API_KEY = os.getenv(\"ELEVENLABS_API_KEY\")\nBASE_URL = \"https://api.elevenlabs.io/v1\"\n\ndef text_to_speech(text: str, voice_id: str = \"EXAVITQu4vr4xnSDxMaL\") -> str:\n    \"\"\"\n    Sends `text` to ElevenLabs and returns a temporary URL to the generated audio.\n    \"\"\"\n    url = f\"{BASE_URL}/text-to-speech/{voice_id}\"\n    headers = {\n        \"xi-api-key\": ELEVEN_API_KEY,\n        \"Content-Type\": \"application/json\",\n    }\n    payload = {\n        \"text\": text,\n        \"model_id\": \"eleven_monolingual_v1\",\n        \"voice_settings\": {\n            \"stability\": 0.75,\n            \"similarity_boost\": 0.85\n        }\n    }\n\n    response = requests.post(url, headers=headers, json=payload)\n    response.raise_for_status()\n    # The API returns the raw audio bytes; we’ll write them to a file.\n    return response.content\n```\n\n`curl`\nIf you prefer a quick test from the command line, here’s the equivalent `curl` call:\n\n```\ncurl -X POST \"https://api.elevenlabs.io/v1/text-to-speech/EXAVITQu4vr4xnSDxMaL\" \\\n  -H \"xi-api-key: $ELEVENLABS_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n        \"text\": \"Welcome to the AI Podcast Generator tutorial.\",\n        \"model_id\": \"eleven_monolingual_v1\",\n        \"voice_settings\": { \"stability\": 0.75, \"similarity_boost\": 0.85 }\n      }' \\\n  --output episode_intro.mp3\n```\n\nReplace `EXAVITQu4vr4xnSDxMaL` with the voice ID you want (ElevenLabs provides a few default voices; you can also upload a custom clone).\n\nMost podcasts are longer than a single API call can comfortably handle (the API caps at ~5 k characters per request). The typical approach is to split the script into logical sections (intro, interview, outro) and generate each chunk separately. Below is a simple orchestrator that does exactly that:\n\n``` python\n# generate_episode.py\nimport os\nimport json\nimport subprocess\nfrom pathlib import Path\nfrom eleven import text_to_speech\n\n# ----------------------------------------------------------------------\n# 1️⃣ Load your script (plain text, one paragraph per line)\n# ----------------------------------------------------------------------\nSCRIPT_PATH = Path(\"script.txt\")\nsegments = SCRIPT_PATH.read_text(encoding=\"utf-8\").split(\"\\n\\n\")  # double newline = segment\n\n# ----------------------------------------------------------------------\n# 2️⃣ Generate audio for each segment\n# ----------------------------------------------------------------------\naudio_files = []\nfor i, segment in enumerate(segments, start=1):\n    print(f\"Generating segment {i}/{len(segments)} …\")\n    audio_bytes = text_to_speech(segment.strip())\n    out_path = Path(f\"segment_{i:03}.mp3\")\n    out_path.write_bytes(audio_bytes)\n    audio_files.append(str(out_path))\n\n# ----------------------------------------------------------------------\n# 3️⃣ Concatenate with ffmpeg\n# ----------------------------------------------------------------------\nlist_file = \"concat_list.txt\"\nwith open(list_file, \"w\") as f:\n    for fp in audio_files:\n        f.write(f\"file '{fp}'\\n\")\n\nfinal_mp3 = \"episode_full.mp3\"\nsubprocess.run(\n    [\"ffmpeg\", \"-y\", \"-f\", \"concat\", \"-safe\", \"0\", \"-i\", list_file,\n     \"-c\", \"copy\", final_mp3],\n    check=True\n)\n\nprint(f\"\\n✅ Episode assembled: {final_mp3}\")\n```\n\n`script.txt`) is split on double newlines. Feel free to adjust the delimiter to match your writing style.\n`text_to_speech`. The function returns raw MP3 bytes, which we save to `segment_###.mp3`.\n`ffmpeg` reads a tiny manifest (`concat_list.txt`) and concatenates the files without re‑encoding, preserving the original quality.\nA podcast isn’t just a voice; you probably want intro music, background ambience, or a short ad slot. The same `ffmpeg` concat method can merge any number of audio tracks. Here’s a quick example that adds a 5‑second intro music clip:\n\n```\n# Prepare a manifest that interleaves music and voice\ncat > concat_list.txt <<EOF\nfile 'intro_music.mp3'\nfile 'segment_001.mp3'\nfile 'segment_002.mp3'\n# …\nfile 'outro_music.mp3'\nEOF\n\nffmpeg -y -f concat -safe 0 -i concat_list.txt -c copy final_podcast.mp3\n```\n\nMake sure your music files have the same sample rate and channel layout as the TTS output (44.1 kHz, stereo) to avoid re‑encoding.\n\nYou can run the script locally, but for a production‑grade pipeline you’ll likely want a serverless function or a small Flask API. Below is a minimal Flask wrapper that accepts a JSON payload with a `script` field and returns a signed URL to the generated episode (using a temporary S3 bucket, for example). The core logic stays the same—just move the code from `generate_episode.py` into a function.\n\n``` python\n# app.py\nfrom flask import Flask, request, jsonify\nfrom eleven import text_to_speech\nimport boto3, os, uuid, subprocess\n\napp = Flask(__name__)\ns3 = boto3.client(\"s3\")\nBUCKET = os.getenv(\"S3_BUCKET\")\n\ndef build_episode(script: str) -> str:\n    segments = script.split(\"\\n\\n\")\n    tmp_dir = f\"/tmp/{uuid.uuid4()}\"\n    os.makedirs(tmp_dir, exist_ok=True)\n\n    audio_paths = []\n    for i, seg in enumerate(segments, 1):\n        audio = text_to_speech(seg.strip())\n        path = f\"{tmp_dir}/seg_{i:03}.mp3\"\n        open(path, \"wb\").write(audio)\n        audio_paths.append(path)\n\n    list_file = f\"{tmp_dir}/list.txt\"\n    with open(list_file, \"w\") as f:\n        for p in audio_paths:\n            f.write(f\"file '{p}'\\n\")\n\n    final_path = f\"{tmp_dir}/episode.mp3\"\n    subprocess.run(\n        [\"ffmpeg\", \"-y\", \"-f\", \"concat\", \"-safe\", \"0\", \"-i\", list_file,\n         \"-c\", \"copy\", final_path],\n        check=True\n    )\n\n    key = f\"episodes/{uuid.uuid4()}.mp3\"\n    s3.upload_file(final_path, BUCKET, key, ExtraArgs={\"ACL\": \"public-read\"})\n    return f\"https://{BUCKET}.s3.amazonaws.com/{key}\"\n\n@app.route(\"/generate\", methods=[\"POST\"])\ndef generate():\n    data = request.get_json()\n    if not data or \"script\" not in data:\n        return jsonify({\"error\": \"Missing script\"}), 400\n    url = build_episode(data[\"script\"])\n    return jsonify({\"episode_url\": url})\n\nif __name__ == \"__main__\":\n    app.run(debug=True)\n```\n\nDeploy this to a platform like **Render**, **Fly.io**, or **AWS Lambda + API Gateway** and you have a fully automated podcast generator that anyone can call via a simple HTTP request.\n\n| Tip | Why it matters | \n|---|---|\n| **Keep sentences under 150 characters** | ElevenLabs handles short bursts more naturally; longer sentences can produce slight breath artifacts. | \n| **Add pauses with `,` or `...`** | The engine interprets punctuation as breath or pause cues, giving a more human rhythm. | \n| **Use the same voice ID for every episode** | Consistency builds brand identity. Clone your own voice if you want a unique host. | \n| **Test stability & similarity settings** | Higher stability yields smoother speech; similarity boost makes the voice sound more like the reference. | \n\nFeel free to experiment with the `voice_settings` payload – the API docs (linked from the ElevenLabs dashboard) provide a nice interactive playground.\n\nYou now have a complete end‑to‑end workflow:\n\nAll of this is powered by **ElevenLabs**, whose realistic voice cloning makes the final product sound like a professional narrator rather than a robotic read‑out.  \n\nGive it a spin, tweak the voice settings, and start churning out episodes without ever stepping into a recording booth.\n\n**Ready to give your podcast a voice?** Sign up at [https://try.elevenlabs.io/kr07zfuqn1bp](https://try.elevenlabs.io/kr07zfuqn1bp) and start generating today!", "url": "https://wpnews.pro/news/create-an-ai-podcast-generator-with-elevenlabs", "canonical_source": "https://dev.to/voice_developer/create-an-ai-podcast-generator-with-elevenlabs-4e9m", "published_at": "2026-09-28 00:19:53+00:00", "updated_at": "2026-09-28 00:30:47.641162+00:00", "lang": "en", "topics": ["ai-tools", "generative-ai", "ai-products", "developer-tools"], "entities": ["ElevenLabs", "Python", "ffmpeg", "curl"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/create-an-ai-podcast-generator-with-elevenlabs", "markdown": "https://wpnews.pro/news/create-an-ai-podcast-generator-with-elevenlabs.md", "text": "https://wpnews.pro/news/create-an-ai-podcast-generator-with-elevenlabs.txt", "jsonld": "https://wpnews.pro/news/create-an-ai-podcast-generator-with-elevenlabs.jsonld"}}