Chatbots have become the go‑to interface for customer support, personal assistants, and even hobby projects. Adding a voice layer turns a static text bot into a more natural, hands‑free experience. With the rise of powerful large‑language models (LLMs) and high‑quality text‑to‑speech (TTS) services, you can spin up a voice‑first assistant in a single afternoon.
In this guide we’ll stitch together OpenAI’s GPT‑4 for conversational intelligence and ElevenLabs for realistic speech synthesis. By the end you’ll have a Python script that listens to your microphone, sends the transcript to OpenAI, and speaks the response back using ElevenLabs’ voice cloning technology.
Tip: If you’re looking for a quick way to get lifelike speech, check out ElevenLabs here: https://try.elevenlabs.io/kr07zfuqn1bp
| Item | Reason |
|---|---|
| Python 3.9+ | Core language for the demo |
openai Python package |
Calls the GPT‑4 API |
pyaudio orsounddevice |
Capture microphone audio |
requests |
Send HTTP requests to ElevenLabs |
| ElevenLabs API key | Access to their TTS endpoints |
| OpenAI API key | Talk to GPT‑4 |
You can install the required packages with:
pip install openai requests sounddevice numpy scipy
(If you prefer pyaudio, replace sounddevice with pyaudio.)
First, grab an API key from the OpenAI dashboard and store it securely, e.g. in an environment variable:
export OPENAI_API_KEY="sk-..."
A tiny helper function to query GPT‑4 looks like this:
import os
import openai
openai.api_key = os.getenv("OPENAI_API_KEY")
def ask_gpt(prompt: str) -> str:
response = openai.ChatCompletion.create(
model="gpt-4o-mini", # or "gpt-4" if you have access
messages=[{"role": "user", "content": prompt}],
temperature=0.7,
)
return response.choices[0].message["content"].strip()
Feel free to tweak temperature, max_tokens, or the model name to suit your use‑case.
ElevenLabs provides a simple REST endpoint that accepts plain text and returns an MP3 (or WAV) audio stream. Sign up at the affiliate link to obtain an API key: https://try.elevenlabs.io/kr07zfuqn1bp
Here’s a minimal wrapper:
import os
import requests
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
VOICE_ID = "EXAVITQu4vr4xnSDxMaL" # default “Rachel” voice; replace with your cloned voice ID
def synthesize(text: str) -> bytes:
url = f"https://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}"
headers = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json"
}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1", # the high‑quality model
"voice_settings": {
"stability": 0.75,
"similarity_boost": 0.85
}
}
resp = requests.post(url, json=payload, headers=headers)
resp.raise_for_status()
return resp.content # raw audio bytes (MP3)
If you have a custom cloned voice, replace VOICE_ID with the ID you receive after up your voice sample.
Below is a complete script that:
import os, io, time, numpy as np, sounddevice as sd, scipy.io.wavfile as wav
import openai, requests
openai.api_key = os.getenv("OPENAI_API_KEY")
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
VOICE_ID = "EXAVITQu4vr4xnSDxMaL" # change if you have a custom voice
def record(duration=3, fs=16000):
print("🎤 Listening…")
audio = sd.rec(int(duration * fs), samplerate=fs, channels=1, dtype='int16')
sd.wait()
return audio.squeeze()
def transcribe(audio_np):
buf = io.BytesIO()
wav.write(buf, 16000, audio_np)
buf.seek(0)
transcript = openai.Audio.transcribe("whisper-1", buf)
return transcript["text"]
def ask_gpt(prompt):
resp = openai.ChatCompletion.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
temperature=0.7,
)
return resp.choices[0].message["content"].strip()
def synthesize(text):
url = f"https://api.elevenlabs.io/v1/text-to-speech/{VOICE_ID}"
headers = {"xi-api-key": ELEVEN_API_KEY, "Content-Type": "application/json"}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1",
"voice_settings": {"stability": 0.75, "similarity_boost": 0.85}
}
r = requests.post(url, json=payload, headers=headers)
r.raise_for_status()
return r.content
def play(audio_bytes):
import pygame
pygame.mixer.init()
sound = pygame.mixer.Sound(io.BytesIO(audio_bytes))
sound.play()
while pygame.mixer.get_busy():
time.sleep(0.1)
if __name__ == "__main__":
while True:
try:
raw = record()
user_text = transcribe(raw)
print(f"🗣️ You said: {user_text}")
reply = ask_gpt(user_text)
print(f"🤖 Bot: {reply}")
audio = synthesize(reply)
play(audio)
except KeyboardInterrupt:
print("\n👋 Bye!")
break
except Exception as e:
print(f"❗ Error: {e}")
What’s happening under the hood?
If you need lower latency, consider streaming audio to Whisper via the audio.transcriptions endpoint and using ElevenLabs’ streaming TTS (available in their beta). The pattern stays the same—just replace the blocking requests.post with a websocket client.
Wrap the ask_gpt and synthesize calls into an HTTP endpoint (e.g., FastAPI or AWS Lambda). Front‑end apps can then send text or audio payloads and receive a URL to an MP3 that can be streamed directly in the browser.
ElevenLabs shines when you upload a few seconds of your own voice and let the service generate a personalized voice ID. Use the same synthesize function; just swap VOICE_ID with the ID returned after the cloning process. The result feels like you’re talking to yourself!
You now have a fully functional voice‑enabled chatbot built with OpenAI and ElevenLabs. The core ideas—record, transcribe, generate, synthesize—are reusable across many domains: virtual assistants, language learning tools, accessibility apps, and more.
If you enjoyed the demo and want to experiment with higher‑quality voices or custom clones, give ElevenLabs a spin: https://try.elevenlabs.io/kr07zfuqn1bp. Their API is developer‑friendly, fast, and the speech quality is truly impressive. Happy coding!