{"slug": "elevenlabs-twilio-building-voice-call-applications", "title": "ElevenLabs + Twilio: Building Voice Call Applications", "summary": "A developer has published a step-by-step guide for building an AI-powered voice call application that combines Twilio's telephony webhooks with ElevenLabs' text-to-speech API and OpenAI's GPT-4o-mini for conversational responses. The Flask-based example routes inbound calls through Twilio's speech recognition, generates a reply via ChatGPT, and plays back ElevenLabs-generated audio, with ngrok used to expose the local server to Twilio's webhook.", "body_md": "If you’ve ever wanted to turn a simple phone call into an interactive, AI‑powered experience, you’re in the right place. Twilio gives you the plumbing to make and receive voice calls, while ElevenLabs provides state‑of‑the‑art text‑to‑speech (TTS) and voice cloning. Put them together and you can build everything from personalized voicemail assistants to real‑time language translation over the phone.\n\nIn this post we’ll walk through a minimal but fully functional example:\n\n`<Gather>`)\nBy the end you’ll have a reusable Flask endpoint that you can deploy on any cloud provider.\n\n**Pro tip:** If you need a high‑quality voice for your brand, ElevenLabs’ voice cloning lets you upload a few minutes of audio and get a custom voice that sounds like a real person. Check it out here: [https://try.elevenlabs.io/kr07zfuqn1bp](https://try.elevenlabs.io/kr07zfuqn1bp)\n\n| What you need | Why | \n|---|---|\n| **Twilio account** | To get a phone number and expose a webhook | \n| **ElevenLabs API key** | To call the TTS endpoint | \n| **OpenAI API key** (optional) | For speech‑to‑text and conversational AI | \n| **Python 3.9+** | The example uses Flask | \n| **ngrok** (or any public URL) | To expose your local server to Twilio | \n\nInstall the required Python packages:\n\n```\npip install flask twilio requests openai\n```\n\nCreate a new phone number in the Twilio console and point its **Voice & Fax → A CALL COMES IN** webhook to `https://<your‑public‑url>/voice`. When Twilio receives a call it will POST an XML document (Twiml) that tells it what to do next.\n\n``` python\n# app.py\nfrom flask import Flask, request, Response\nfrom twilio.twiml.voice_response import VoiceResponse, Gather\n\napp = Flask(__name__)\n\n@app.route(\"/voice\", methods=[\"POST\"])\ndef voice():\n    resp = VoiceResponse()\n    # Ask the caller to speak after the beep\n    gather = Gather(input=\"speech\", action=\"/process\", method=\"POST\", timeout=5)\n    gather.say(\"Hi! Tell me what you need help with.\")\n    resp.append(gather)\n    # If no speech was captured, fallback\n    resp.say(\"Sorry, I didn't catch that. Goodbye!\")\n    resp.hangup()\n    return Response(str(resp), mimetype=\"application/xml\")\n```\n\nRun the app locally and expose it with ngrok:\n\n```\npython app.py\nngrok http 5000\n```\n\nTake the HTTPS ngrok URL (e.g., `https://abcd1234.ngrok.io`) and paste it into the Twilio webhook field.\n\nWhen the caller finishes speaking, Twilio sends a POST to `/process` with a `SpeechResult` field that already contains a transcription (thanks to Twilio’s built‑in speech recognition). If you want higher accuracy or multiple languages, you can pipe the raw audio to Whisper instead, but for this demo we’ll use Twilio’s result directly.\n\n``` python\nimport os\nimport openai\nimport requests\n\nELEVEN_API_KEY = os.getenv(\"ELEVEN_API_KEY\")\nOPENAI_API_KEY = os.getenv(\"OPENAI_API_KEY\")\nELEVEN_TTS_URL = \"https://api.elevenlabs.io/v1/text-to-speech\"\n\nopenai.api_key = OPENAI_API_KEY\n\n@app.route(\"/process\", methods=[\"POST\"])\ndef process():\n    # Grab what Twilio recognized\n    user_text = request.form.get(\"SpeechResult\", \"\")\n    if not user_text:\n        return fallback_response(\"I didn't hear anything.\")\n\n    # Generate a ChatGPT response\n    reply = chatgpt_reply(user_text)\n\n    # Convert reply to audio with ElevenLabs\n    audio_url = eleven_tts(reply)\n\n    # Build TwiML to play the audio\n    resp = VoiceResponse()\n    resp.play(audio_url)\n    resp.hangup()\n    return Response(str(resp), mimetype=\"application/xml\")\nphp\ndef chatgpt_reply(prompt: str) -> str:\n    completion = openai.ChatCompletion.create(\n        model=\"gpt-4o-mini\",\n        messages=[{\"role\": \"user\", \"content\": prompt}],\n        temperature=0.7,\n    )\n    return completion.choices[0].message.content.strip()\n```\n\nElevenLabs’ API accepts plain text and returns a short‑lived URL to an MP3 file. The endpoint we’ll use is `/v1/text-to-speech/{voice_id}`. For most developers the “default” voice (` EXAVITQu4vr4xnSDxMaL`) works great, but you can replace it with a cloned voice ID if you’ve uploaded your own samples.\n\n``` php\ndef eleven_tts(text: str) -> str:\n    voice_id = \"EXAVITQu4vr4xnSDxMaL\"  # default voice\n    headers = {\n        \"xi-api-key\": ELEVEN_API_KEY,\n        \"Content-Type\": \"application/json\"\n    }\n    payload = {\n        \"text\": text,\n        \"model_id\": \"eleven_monolingual_v1\",\n        \"voice_settings\": {\n            \"stability\": 0.75,\n            \"similarity_boost\": 0.85\n        }\n    }\n\n    response = requests.post(\n        f\"{ELEVEN_TTS_URL}/{voice_id}\",\n        json=payload,\n        headers=headers,\n        timeout=15,\n    )\n    response.raise_for_status()\n    # The API returns raw audio bytes; we upload to a temporary storage service\n    # For simplicity, we’ll use ElevenLabs’ own temporary URL feature:\n    return response.json()[\"audio_url\"]\n```\n\n**Note:** If you prefer not to host the MP3 yourself, ElevenLabs can give you a temporary URL that Twilio can stream directly, as shown above.\n\nYour final `app.py` should look roughly like this:\n\n``` python\nimport os\nfrom flask import Flask, request, Response\nfrom twilio.twiml.voice_response import VoiceResponse, Gather\nimport openai\nimport requests\n\napp = Flask(__name__)\n\nELEVEN_API_KEY = os.getenv(\"ELEVEN_API_KEY\")\nOPENAI_API_KEY = os.getenv(\"OPENAI_API_KEY\")\nELEVEN_TTS_URL = \"https://api.elevenlabs.io/v1/text-to-speech\"\nopenai.api_key = OPENAI_API_KEY\n\ndef chatgpt_reply(prompt: str) -> str:\n    completion = openai.ChatCompletion.create(\n        model=\"gpt-4o-mini\",\n        messages=[{\"role\": \"user\", \"content\": prompt}],\n        temperature=0.7,\n    )\n    return completion.choices[0].message.content.strip()\n\ndef eleven_tts(text: str) -> str:\n    voice_id = \"EXAVITQu4vr4xnSDxMaL\"\n    headers = {\n        \"xi-api-key\": ELEVEN_API_KEY,\n        \"Content-Type\": \"application/json\"\n    }\n    payload = {\n        \"text\": text,\n        \"model_id\": \"eleven_monolingual_v1\",\n        \"voice_settings\": {\"stability\": 0.75, \"similarity_boost\": 0.85}\n    }\n    resp = requests.post(f\"{ELEVEN_TTS_URL}/{voice_id}\", json=payload, headers=headers)\n    resp.raise_for_status()\n    return resp.json()[\"audio_url\"]\n\ndef fallback_response(message: str):\n    resp = VoiceResponse()\n    resp.say(message)\n    resp.hangup()\n    return Response(str(resp), mimetype=\"application/xml\")\n\n@app.route(\"/voice\", methods=[\"POST\"])\ndef voice():\n    resp = VoiceResponse()\n    gather = Gather(input=\"speech\", action=\"/process\", method=\"POST\", timeout=5)\n    gather.say(\"Hey there! What can I help you with today?\")\n    resp.append(gather)\n    resp.say(\"Sorry, I didn't hear anything. Bye!\")\n    resp.hangup()\n    return Response(str(resp), mimetype=\"application/xml\")\n\n@app.route(\"/process\", methods=[\"POST\"])\ndef process():\n    user_text = request.form.get(\"SpeechResult\", \"\")\n    if not user_text:\n        return fallback_response(\"I didn't catch that.\")\n\n    reply = chatgpt_reply(user_text)\n    audio_url = eleven_tts(reply)\n\n    resp = VoiceResponse()\n    resp.play(audio_url)\n    resp.hangup()\n    return Response(str(resp), mimetype=\"application/xml\")\n\nif __name__ == \"__main__\":\n    app.run(debug=True, port=5000)\n```\n\nDeploy the script to your favorite host (Heroku, Fly.io, Render, etc.), point the Twilio webhook at the live URL, and you’re ready to make calls that sound like a real person—powered by ElevenLabs.\n\n| Feature | How to add it | \n|---|---|\n| **Voice cloning** | Upload a few minutes of your own voice to ElevenLabs, grab the returned `voice_id` , and replace the default`voice_id` in`eleven_tts` . | \n| **Multi‑language support** | Use Whisper or Azure Speech‑to‑Text for transcription, then set `model_id` to`eleven_multilingual_v2` when calling ElevenLabs. | \n| **Persisted conversation** | Store the `conversation_id` from OpenAI and pass it back on each turn to maintain context. | \n| **Call recording** | Add `<Record>` in the TwiML to capture the whole interaction for compliance or analytics. | \n| **Interactive menus** | Use multiple `<Gather>` blocks and DTMF (`input=\"dtmf\"` ) to build IVR trees. | \n\nBuilding a voice‑first AI app used to require a deep dive into DSP libraries and self‑hosted TTS engines. With Twilio handling the telephony plumbing and ElevenLabs delivering crystal‑clear, human‑like speech, the barrier to entry is now a few lines of code.\n\nReady to give your callers a voice that sounds truly human? Grab an API key from ElevenLabs and start experimenting today:\n\n**👉 Try ElevenLabs now: [https://try.elevenlabs.io/kr07zfuqn1bp](https://try.elevenlabs.io/kr07zfuqn1bp)**\n\nHappy coding, and may your next voice app sound as good as it feels!", "url": "https://wpnews.pro/news/elevenlabs-twilio-building-voice-call-applications", "canonical_source": "https://dev.to/voice_developer/elevenlabs-twilio-building-voice-call-applications-i0k", "published_at": "2026-10-01 23:00:21+00:00", "updated_at": "2026-10-01 23:14:36.561494+00:00", "lang": "en", "topics": ["ai-products", "ai-tools", "natural-language-processing", "developer-tools"], "entities": ["ElevenLabs", "Twilio", "OpenAI", "Flask", "ngrok", "GPT-4o-mini", "Whisper"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/elevenlabs-twilio-building-voice-call-applications", "markdown": "https://wpnews.pro/news/elevenlabs-twilio-building-voice-call-applications.md", "text": "https://wpnews.pro/news/elevenlabs-twilio-building-voice-call-applications.txt", "jsonld": "https://wpnews.pro/news/elevenlabs-twilio-building-voice-call-applications.jsonld"}}