{"slug": "build-a-voice-enabled-chatbot-with-elevenlabs-and-openai", "title": "Build a Voice-Enabled Chatbot with ElevenLabs and OpenAI", "summary": "A developer published a step-by-step guide for building a voice-enabled chatbot that chains OpenAI's GPT-4 and Whisper with ElevenLabs' text-to-speech API. The Python script records microphone audio, transcribes it via Whisper, sends the transcript to GPT-4 for a response, and synthesizes the reply into speech using ElevenLabs' voice cloning endpoints.", "body_md": "Chatbots have become the go‑to interface for customer support, personal assistants, and even hobby projects. Adding a voice layer turns a static text bot into a more natural, hands‑free experience. With the rise of powerful large‑language models (LLMs) and high‑quality text‑to‑speech (TTS) services, you can spin up a voice‑first assistant in a single afternoon.\n\nIn this guide we’ll stitch together **OpenAI’s GPT‑4** for conversational intelligence and **ElevenLabs** for realistic speech synthesis. By the end you’ll have a Python script that listens to your microphone, sends the transcript to OpenAI, and speaks the response back using ElevenLabs’ voice cloning technology.\n\n**Tip:** If you’re looking for a quick way to get lifelike speech, check out ElevenLabs here: [https://try.elevenlabs.io/kr07zfuqn1bp](https://try.elevenlabs.io/kr07zfuqn1bp)\n\n| Item | Reason | \n|---|---|\n| Python 3.9+ | Core language for the demo | \n| `openai` Python package | Calls the GPT‑4 API | \n| `pyaudio` or`sounddevice` | Capture microphone audio | \n| `requests` | Send HTTP requests to ElevenLabs | \n| ElevenLabs API key | Access to their TTS endpoints | \n| OpenAI API key | Talk to GPT‑4 | \n\nYou can install the required packages with:\n\n```\npip install openai requests sounddevice numpy scipy\n```\n\n(If you prefer `pyaudio`, replace `sounddevice` with `pyaudio`.)\n\nFirst, grab an API key from the OpenAI dashboard and store it securely, e.g. in an environment variable:\n\n```\nexport OPENAI_API_KEY=\"sk-...\"\n```\n\nA tiny helper function to query GPT‑4 looks like this:\n\n``` python\nimport os\nimport openai\n\nopenai.api_key = os.getenv(\"OPENAI_API_KEY\")\n\ndef ask_gpt(prompt: str) -> str:\n    response = openai.ChatCompletion.create(\n        model=\"gpt-4o-mini\",   # or \"gpt-4\" if you have access\n        messages=[{\"role\": \"user\", \"content\": prompt}],\n        temperature=0.7,\n    )\n    return response.choices[0].message[\"content\"].strip()\n```\n\nFeel free to tweak `temperature`, `max_tokens`, or the `model` name to suit your use‑case.\n\nElevenLabs provides a simple REST endpoint that accepts plain text and returns an MP3 (or WAV) audio stream. Sign up at the affiliate link to obtain an API key: [https://try.elevenlabs.io/kr07zfuqn1bp](https://try.elevenlabs.io/kr07zfuqn1bp)\n\nHere’s a minimal wrapper:\n\n``` python\nimport os\nimport requests\n\nELEVEN_API_KEY = os.getenv(\"ELEVEN_API_KEY\")\nVOICE_ID = \"EXAVITQu4vr4xnSDxMaL\"   # default “Rachel” voice; replace with your cloned voice ID\n\ndef synthesize(text: str) -> bytes:\n    url = f\"https://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}\"\n    headers = {\n        \"xi-api-key\": ELEVEN_API_KEY,\n        \"Content-Type\": \"application/json\"\n    }\n    payload = {\n        \"text\": text,\n        \"model_id\": \"eleven_monolingual_v1\",  # the high‑quality model\n        \"voice_settings\": {\n            \"stability\": 0.75,\n            \"similarity_boost\": 0.85\n        }\n    }\n    resp = requests.post(url, json=payload, headers=headers)\n    resp.raise_for_status()\n    return resp.content   # raw audio bytes (MP3)\n```\n\nIf you have a custom cloned voice, replace `VOICE_ID` with the ID you receive after uploading your voice sample.\n\nBelow is a complete script that:\n\n``` python\nimport os, io, time, numpy as np, sounddevice as sd, scipy.io.wavfile as wav\nimport openai, requests\n\n# ----- Config -----\nopenai.api_key = os.getenv(\"OPENAI_API_KEY\")\nELEVEN_API_KEY = os.getenv(\"ELEVEN_API_KEY\")\nVOICE_ID = \"EXAVITQu4vr4xnSDxMaL\"   # change if you have a custom voice\n\n# ----- Helper functions -----\ndef record(duration=3, fs=16000):\n    print(\"🎤 Listening…\")\n    audio = sd.rec(int(duration * fs), samplerate=fs, channels=1, dtype='int16')\n    sd.wait()\n    return audio.squeeze()\n\ndef transcribe(audio_np):\n    # Convert numpy array to WAV bytes for Whisper\n    buf = io.BytesIO()\n    wav.write(buf, 16000, audio_np)\n    buf.seek(0)\n    transcript = openai.Audio.transcribe(\"whisper-1\", buf)\n    return transcript[\"text\"]\n\ndef ask_gpt(prompt):\n    resp = openai.ChatCompletion.create(\n        model=\"gpt-4o-mini\",\n        messages=[{\"role\": \"user\", \"content\": prompt}],\n        temperature=0.7,\n    )\n    return resp.choices[0].message[\"content\"].strip()\n\ndef synthesize(text):\n    url = f\"https://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}\"\n    headers = {\"xi-api-key\": ELEVEN_API_KEY, \"Content-Type\": \"application/json\"}\n    payload = {\n        \"text\": text,\n        \"model_id\": \"eleven_monolingual_v1\",\n        \"voice_settings\": {\"stability\": 0.75, \"similarity_boost\": 0.85}\n    }\n    r = requests.post(url, json=payload, headers=headers)\n    r.raise_for_status()\n    return r.content\n\ndef play(audio_bytes):\n    # ElevenLabs returns MP3; convert to raw PCM for playback\n    import pygame\n    pygame.mixer.init()\n    sound = pygame.mixer.Sound(io.BytesIO(audio_bytes))\n    sound.play()\n    while pygame.mixer.get_busy():\n        time.sleep(0.1)\n\n# ----- Main loop -----\nif __name__ == \"__main__\":\n    while True:\n        try:\n            # 1️⃣ Capture voice\n            raw = record()\n            # 2️⃣ Transcribe\n            user_text = transcribe(raw)\n            print(f\"🗣️ You said: {user_text}\")\n\n            # 3️⃣ Get LLM reply\n            reply = ask_gpt(user_text)\n            print(f\"🤖 Bot: {reply}\")\n\n            # 4️⃣ Convert reply to speech\n            audio = synthesize(reply)\n\n            # 5️⃣ Play back\n            play(audio)\n\n        except KeyboardInterrupt:\n            print(\"\\n👋 Bye!\")\n            break\n        except Exception as e:\n            print(f\"❗ Error: {e}\")\n```\n\n**What’s happening under the hood?**\n\nIf you need lower latency, consider streaming audio to Whisper via the `audio.transcriptions` endpoint and using ElevenLabs’ **streaming TTS** (available in their beta). The pattern stays the same—just replace the blocking `requests.post` with a websocket client.\n\nWrap the `ask_gpt` and `synthesize` calls into an HTTP endpoint (e.g., FastAPI or AWS Lambda). Front‑end apps can then send text or audio payloads and receive a URL to an MP3 that can be streamed directly in the browser.\n\nElevenLabs shines when you upload a few seconds of your own voice and let the service generate a personalized voice ID. Use the same `synthesize` function; just swap `VOICE_ID` with the ID returned after the cloning process. The result feels like you’re talking to yourself!\n\nYou now have a fully functional voice‑enabled chatbot built with OpenAI and ElevenLabs. The core ideas—record, transcribe, generate, synthesize—are reusable across many domains: virtual assistants, language learning tools, accessibility apps, and more.\n\nIf you enjoyed the demo and want to experiment with higher‑quality voices or custom clones, give ElevenLabs a spin: [https://try.elevenlabs.io/kr07zfuqn1bp](https://try.elevenlabs.io/kr07zfuqn1bp). Their API is developer‑friendly, fast, and the speech quality is truly impressive. Happy coding!", "url": "https://wpnews.pro/news/build-a-voice-enabled-chatbot-with-elevenlabs-and-openai", "canonical_source": "https://dev.to/voice_developer/build-a-voice-enabled-chatbot-with-elevenlabs-and-openai-3afp", "published_at": "2026-09-28 23:43:30+00:00", "updated_at": "2026-09-28 23:49:00.528168+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "generative-ai", "natural-language-processing", "ai-products"], "entities": ["OpenAI", "ElevenLabs", "GPT-4", "Whisper", "Python"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/build-a-voice-enabled-chatbot-with-elevenlabs-and-openai", "markdown": "https://wpnews.pro/news/build-a-voice-enabled-chatbot-with-elevenlabs-and-openai.md", "text": "https://wpnews.pro/news/build-a-voice-enabled-chatbot-with-elevenlabs-and-openai.txt", "jsonld": "https://wpnews.pro/news/build-a-voice-enabled-chatbot-with-elevenlabs-and-openai.jsonld"}}