Podcasting is exploding, but producing high‑quality audio still takes a lot of time and equipment. With modern text‑to‑speech (TTS) and voice‑cloning services you can turn a markdown script into a fully‑produced episode in minutes. In this guide we’ll walk through:
By the end you’ll have a tiny web service that accepts a podcast script (or a URL to a markdown file) and returns a downloadable MP3 ready for publishing.
| Item | Reason |
|---|---|
| ElevenLabs API key – sign up athttps://try.elevenlabs.io/kr07zfuqn1bp | Gives you access to state‑of‑the‑art TTS and voice cloning. |
| Python 3.9+ | Core language for the demo. |
| Flask (or FastAPI) | Minimal web framework to expose a REST endpoint. |
| FFmpeg | Concatenates audio segments and normalises volume. |
| Bluehost account – get one athttps://bluehost.sjv.io/5k0d52 | Provides the cheap, easy‑to‑use hosting environment for our app. |
| Git (optional) | For version control and quick deployment. |
ElevenLabs offers a straightforward REST API. Below is a minimal Python wrapper that sends a text payload and streams back an MP3 file.
import os
import requests
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
BASE_URL = "https://api.elevenlabs.io/v1"
def generate_speech(text: str, voice_id: str = "default") -> bytes:
"""
Calls ElevenLabs TTS and returns raw MP3 bytes.
"""
endpoint = f"{BASE_URL}/text-to-speech/{voice_id}"
headers = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json",
"Accept": "audio/mpeg"
}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1",
"voice_settings": {"stability": 0.75, "similarity_boost": 0.85}
}
response = requests.post(endpoint, json=payload, headers=headers)
response.raise_for_status()
return response.content
Tip: If you want a custom cloned voice, create it in the ElevenLabs dashboard and replace voice_id with the ID you receive.
You can test the function locally:
export ELEVEN_API_KEY=your_key_here
python - <<'PY'
from __main__ import generate_speech
audio = generate_speech("Hello, this is a test of ElevenLabs text‑to‑speech.")
open("test.mp3", "wb").write(audio)
print("Saved test.mp3")
PY
Play test.mp3 to hear the result. The quality is comparable to a human narrator, and the API is priced per character, making it cheap for short‑form podcasts.
Our Flask app will accept a JSON payload like:
{
"title": "AI Podcast 101",
"script_url": "https://example.com/episode1.md",
"intro_music_url": "https://example.com/intro.mp3",
"outro_music_url": "https://example.com/outro.mp3"
}
The flow:
markdown2).
generate_speech for the main body.
import os, subprocess, tempfile, requests, markdown2
from flask import Flask, request, send_file, jsonify
app = Flask(__name__)
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
if not ELEVEN_API_KEY:
raise RuntimeError("Set ELEVEN_API_KEY env var")
from __main__ import generate_speech # assume same file for brevity
def download(url: str) -> bytes:
r = requests.get(url)
r.raise_for_status()
return r.content
def concat_audios(parts: list[bytes]) -> bytes:
"""
Takes a list of raw MP3 bytes, writes them to temp files,
and returns a single concatenated MP3 using ffmpeg.
"""
with tempfile.TemporaryDirectory() as td:
paths = []
for i, data in enumerate(parts):
p = os.path.join(td, f"part{i}.mp3")
open(p, "wb").write(data)
paths.append(p)
list_file = os.path.join(td, "list.txt")
open(list_file, "w").write("\n".join([f"file '{p}'" for p in paths]))
out_path = os.path.join(td, "out.mp3")
cmd = [
"ffmpeg", "-y", "-f", "concat", "-safe", "0",
"-i", list_file, "-c", "copy", out_path
]
subprocess.run(cmd, check=True, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)
return open(out_path, "rb").read()
@app.route("/generate", methods=["POST"])
def generate():
data = request.json
script_md = download(data["script_url"]).decode()
script_text = markdown2.markdown(script_md, strip=True) # strip HTML tags
intro = download(data["intro_music_url"])
narration = generate_speech(script_text)
outro = download(data["outro_music_url"])
final_mp3 = concat_audios([intro, narration, outro])
return send_file(
io.BytesIO(final_mp3),
mimetype="audio/mpeg",
as_attachment=True,
download_name=f"{data['title']}.mp3"
)
if __name__ == "__main__":
app.run(host="0.0.0.0", port=5000)
Dependencies
pip install flask requests markdown2
Run locally with python app.py and POST a JSON payload to http://localhost:5000/generate. You’ll receive a ready‑to‑publish MP3.
requirements.txt, and maps a URL to a WSGI entry point.
All of this can be done without touching a VPS console, making Bluehost the easy, affordable way to spin up a voice‑AI service.
~/my-podcast.
app.py.
flask_app.wsgi (Bluehost will generate it).
Flask
requests
markdown2
pip install -r requirements.txt
ffmpeg via the package manager. If not, you can compile a static binary and place it in your app folder; the Flask code calls
echo "export ELEVEN_API_KEY=your_key_here" >> ~/.bashrc
source ~/.bashrc
https://yourdomain.com/generate.
curl -X POST https://yourdomain.com/generate \
-H "Content-Type: application/json" \
-d '{
"title": "AI Podcast 101",
"script_url": "https://raw.githubusercontent.com/yourrepo/episode1.md",
"intro_music_url": "https://example.com/intro.mp3",
"outro_music_url": "https://example.com/outro.mp3"
}' \
--output episode1.mp3
If the download succeeds, you’ve just built a fully‑hosted AI podcast generator with a few lines of code!
/voices endpoint to upload a few minutes of your own voice and generate a custom voice_id.
All of these extensions sit on the same Flask base; you only need to add a few more routes or background workers.
You now have a practical, production‑ready pipeline:
Give it a spin, tweak the voice settings, and start publishing AI‑generated episodes without ever stepping into a recording studio.
Ready to build? Grab your ElevenLabs API key and a Bluehost account today, and watch your podcast ideas come to life in seconds. Happy coding!