If you’ve ever wanted to turn a simple phone call into an interactive, AI‑powered experience, you’re in the right place. Twilio gives you the plumbing to make and receive voice calls, while ElevenLabs provides state‑of‑the‑art text‑to‑speech (TTS) and voice cloning. Put them together and you can build everything from personalized voicemail assistants to real‑time language translation over the phone.
In this post we’ll walk through a minimal but fully functional example:
<Gather>)
By the end you’ll have a reusable Flask endpoint that you can deploy on any cloud provider.
Pro tip: If you need a high‑quality voice for your brand, ElevenLabs’ voice cloning lets you upload a few minutes of audio and get a custom voice that sounds like a real person. Check it out here: https://try.elevenlabs.io/kr07zfuqn1bp
| What you need | Why |
|---|---|
| Twilio account | To get a phone number and expose a webhook |
| ElevenLabs API key | To call the TTS endpoint |
| OpenAI API key (optional) | For speech‑to‑text and conversational AI |
| Python 3.9+ | The example uses Flask |
| ngrok (or any public URL) | To expose your local server to Twilio |
Install the required Python packages:
pip install flask twilio requests openai
Create a new phone number in the Twilio console and point its Voice & Fax → A CALL COMES IN webhook to https://<your‑public‑url>/voice. When Twilio receives a call it will POST an XML document (Twiml) that tells it what to do next.
from flask import Flask, request, Response
from twilio.twiml.voice_response import VoiceResponse, Gather
app = Flask(__name__)
@app.route("/voice", methods=["POST"])
def voice():
resp = VoiceResponse()
gather = Gather(input="speech", action="/process", method="POST", timeout=5)
gather.say("Hi! Tell me what you need help with.")
resp.append(gather)
resp.say("Sorry, I didn't catch that. Goodbye!")
resp.hangup()
return Response(str(resp), mimetype="application/xml")
Run the app locally and expose it with ngrok:
python app.py
ngrok http 5000
Take the HTTPS ngrok URL (e.g., https://abcd1234.ngrok.io) and paste it into the Twilio webhook field.
When the caller finishes speaking, Twilio sends a POST to /process with a SpeechResult field that already contains a transcription (thanks to Twilio’s built‑in speech recognition). If you want higher accuracy or multiple languages, you can pipe the raw audio to Whisper instead, but for this demo we’ll use Twilio’s result directly.
import os
import openai
import requests
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
ELEVEN_TTS_URL = "https://api.elevenlabs.io/v1/text-to-speech"
openai.api_key = OPENAI_API_KEY
@app.route("/process", methods=["POST"])
def process():
user_text = request.form.get("SpeechResult", "")
if not user_text:
return fallback_response("I didn't hear anything.")
reply = chatgpt_reply(user_text)
audio_url = eleven_tts(reply)
resp = VoiceResponse()
resp.play(audio_url)
resp.hangup()
return Response(str(resp), mimetype="application/xml")
php
def chatgpt_reply(prompt: str) -> str:
completion = openai.ChatCompletion.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
temperature=0.7,
)
return completion.choices[0].message.content.strip()
ElevenLabs’ API accepts plain text and returns a short‑lived URL to an MP3 file. The endpoint we’ll use is /v1/text-to-speech/{voice_id}. For most developers the “default” voice ( EXAVITQu4vr4xnSDxMaL) works great, but you can replace it with a cloned voice ID if you’ve uploaded your own samples.
def eleven_tts(text: str) -> str:
voice_id = "EXAVITQu4vr4xnSDxMaL" # default voice
headers = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json"
}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1",
"voice_settings": {
"stability": 0.75,
"similarity_boost": 0.85
}
}
response = requests.post(
f"{ELEVEN_TTS_URL}/{voice_id}",
json=payload,
headers=headers,
timeout=15,
)
response.raise_for_status()
return response.json()["audio_url"]
Note: If you prefer not to host the MP3 yourself, ElevenLabs can give you a temporary URL that Twilio can stream directly, as shown above.
Your final app.py should look roughly like this:
import os
from flask import Flask, request, Response
from twilio.twiml.voice_response import VoiceResponse, Gather
import openai
import requests
app = Flask(__name__)
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
ELEVEN_TTS_URL = "https://api.elevenlabs.io/v1/text-to-speech"
openai.api_key = OPENAI_API_KEY
def chatgpt_reply(prompt: str) -> str:
completion = openai.ChatCompletion.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
temperature=0.7,
)
return completion.choices[0].message.content.strip()
def eleven_tts(text: str) -> str:
voice_id = "EXAVITQu4vr4xnSDxMaL"
headers = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json"
}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1",
"voice_settings": {"stability": 0.75, "similarity_boost": 0.85}
}
resp = requests.post(f"{ELEVEN_TTS_URL}/{voice_id}", json=payload, headers=headers)
resp.raise_for_status()
return resp.json()["audio_url"]
def fallback_response(message: str):
resp = VoiceResponse()
resp.say(message)
resp.hangup()
return Response(str(resp), mimetype="application/xml")
@app.route("/voice", methods=["POST"])
def voice():
resp = VoiceResponse()
gather = Gather(input="speech", action="/process", method="POST", timeout=5)
gather.say("Hey there! What can I help you with today?")
resp.append(gather)
resp.say("Sorry, I didn't hear anything. Bye!")
resp.hangup()
return Response(str(resp), mimetype="application/xml")
@app.route("/process", methods=["POST"])
def process():
user_text = request.form.get("SpeechResult", "")
if not user_text:
return fallback_response("I didn't catch that.")
reply = chatgpt_reply(user_text)
audio_url = eleven_tts(reply)
resp = VoiceResponse()
resp.play(audio_url)
resp.hangup()
return Response(str(resp), mimetype="application/xml")
if __name__ == "__main__":
app.run(debug=True, port=5000)
Deploy the script to your favorite host (Heroku, Fly.io, Render, etc.), point the Twilio webhook at the live URL, and you’re ready to make calls that sound like a real person—powered by ElevenLabs.
| Feature | How to add it |
|---|---|
| Voice cloning | Upload a few minutes of your own voice to ElevenLabs, grab the returned voice_id , and replace the defaultvoice_id ineleven_tts . |
| Multi‑language support | Use Whisper or Azure Speech‑to‑Text for transcription, then set model_id toeleven_multilingual_v2 when calling ElevenLabs. |
| Persisted conversation | Store the conversation_id from OpenAI and pass it back on each turn to maintain context. |
| Call recording | Add <Record> in the TwiML to capture the whole interaction for compliance or analytics. |
| Interactive menus | Use multiple <Gather> blocks and DTMF (input="dtmf" ) to build IVR trees. |
Building a voice‑first AI app used to require a deep dive into DSP libraries and self‑hosted TTS engines. With Twilio handling the telephony plumbing and ElevenLabs delivering crystal‑clear, human‑like speech, the barrier to entry is now a few lines of code.
Ready to give your callers a voice that sounds truly human? Grab an API key from ElevenLabs and start experimenting today:
👉 Try ElevenLabs now: https://try.elevenlabs.io/kr07zfuqn1bp
Happy coding, and may your next voice app sound as good as it feels!