Adding AI Voice to Discord Bots A developer published a step-by-step guide for adding AI text-to-speech to Discord bots using ElevenLabs' REST API and Python's discord.py library. The tutorial wraps ElevenLabs' text-to-speech endpoint in an async aiohttp call, pipes the returned MP3 bytes through FFmpeg into Opus format, and plays the audio in a voice channel via discord.FFmpegPCMAudio, enabling bots to speak messages or announcements in pre-built or cloned voices. Discord has become the go‑to place for communities, gaming squads, study groups, and even professional meet‑ups. Most bots you’ll find on the platform are text‑only, but adding a voice component can dramatically boost engagement: In the past, developers had to host their own TTS engines or rely on Discord’s built‑in “Speak” feature, which is limited to raw audio streams. Today, services like ElevenLabs make high‑quality text‑to‑speech TTS and voice cloning a breeze, and they expose simple HTTP APIs that you can call from any language. Below you’ll find a step‑by‑step guide to hooking up ElevenLabs to a Discord bot using Python and the discord.py library. By the end you’ll have a bot that can read messages, announce events, or even speak in a cloned voice that matches your brand. pip install -U discord.py . apt-get install ffmpeg , brew install ffmpeg . ElevenLabs provides a straightforward REST endpoint. We’ll create a tiny wrapper that sends plain text and receives an MP3 file. python import aiohttp import os ELEVEN API KEY = os.getenv "ELEVEN API KEY" ELEVEN TTS URL = "https://api.elevenlabs.io/v1/text-to-speech" async def synthesize text: str, voice id: str = "EXAVITQu4vr4xnSDxMaL" - bytes: """ Convert text to speech using ElevenLabs. Returns raw MP3 bytes. """ headers = { "xi-api-key": ELEVEN API KEY, "Content-Type": "application/json", } payload = { "text": text, "voice settings": {"stability": 0.75, "similarity boost": 0.85}, } async with aiohttp.ClientSession as session: async with session.post f"{ELEVEN TTS URL}/{voice id}", json=payload, headers=headers, as resp: resp.raise for status return await resp.read The voice id can be any of ElevenLabs’ pre‑built voices or a custom clone you created in the dashboard. Discord expects audio in Opus format inside an AudioSource . The easiest way is to pipe the MP3 through FFmpeg and let discord.py handle the streaming. php import discord import io import subprocess def mp3 to opus mp3 bytes: bytes - discord.FFmpegPCMAudio: """ Takes raw MP3 bytes, writes them to a pipe, and returns an FFmpegPCMAudio source. """ Use a BytesIO object as a temporary file mp3 buffer = io.BytesIO mp3 bytes FFmpeg command – reads from stdin -i pipe:0 and outputs opus to stdout ffmpeg options = { "before options": "-nostdin", "options": "-f s16le -ar 48000 -ac 2 pipe:1", } return discord.FFmpegPCMAudio mp3 buffer, ffmpeg options Below is a minimal bot that joins a voice channel when you type join , then reads any subsequent message that starts with say . python import discord from discord.ext import commands intents = discord.Intents.default intents.message content = True Needed for reading message text bot = commands.Bot command prefix=" ", intents=intents @bot.event async def on ready : print f"🤖 {bot.user} is ready " @bot.command async def join ctx : """Bot joins the author's voice channel.""" if ctx.author.voice: channel = ctx.author.voice.channel await channel.connect await ctx.send f"Joined {channel.name} 🎤" else: await ctx.send "You need to be in a voice channel first " @bot.command async def say ctx, , message: str : """Bot speaks the supplied text using ElevenLabs.""" if not ctx.voice client: await ctx.send "I'm not in a voice channel. Use join first." return 1️⃣ Synthesize speech try: mp3 data = await synthesize message except Exception as e: await ctx.send f"❗ TTS error: {e}" return 2️⃣ Convert to Opus and play audio source = mp3 to opus mp3 data ctx.voice client.play audio source, after=lambda e: print "Finished playing", e await ctx.send f"🗣 Speaking: {message} " @bot.command async def leave ctx : """Disconnects the bot from voice.""" if ctx.voice client: await ctx.voice client.disconnect await ctx.send "Goodbye 👋" else: await ctx.send "I'm not connected to any voice channel." bot.run os.getenv "DISCORD TOKEN" join say