# Adding AI Voice to Discord Bots

> Source: <https://dev.to/voice_developer/adding-ai-voice-to-discord-bots-pof>
> Published: 2026-10-01 01:06:32+00:00

Discord has become the go‑to place for communities, gaming squads, study groups, and even professional meet‑ups. Most bots you’ll find on the platform are text‑only, but adding a voice component can dramatically boost engagement:

In the past, developers had to host their own TTS engines or rely on Discord’s built‑in “Speak” feature, which is limited to raw audio streams. Today, services like **ElevenLabs** make high‑quality text‑to‑speech (TTS) and voice cloning a breeze, and they expose simple HTTP APIs that you can call from any language.

Below you’ll find a step‑by‑step guide to hooking up ElevenLabs to a Discord bot using Python and the `discord.py` library. By the end you’ll have a bot that can read messages, announce events, or even speak in a cloned voice that matches your brand.

`pip install -U discord.py`.
`apt-get install ffmpeg`, `brew install ffmpeg`).
ElevenLabs provides a straightforward REST endpoint. We’ll create a tiny wrapper that sends plain text and receives an MP3 file.

``` python
import aiohttp
import os

ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
ELEVEN_TTS_URL = "https://api.elevenlabs.io/v1/text-to-speech"

async def synthesize(text: str, voice_id: str = "EXAVITQu4vr4xnSDxMaL") -> bytes:
    """
    Convert `text` to speech using ElevenLabs.
    Returns raw MP3 bytes.
    """
    headers = {
        "xi-api-key": ELEVEN_API_KEY,
        "Content-Type": "application/json",
    }
    payload = {
        "text": text,
        "voice_settings": {"stability": 0.75, "similarity_boost": 0.85},
    }
    async with aiohttp.ClientSession() as session:
        async with session.post(
            f"{ELEVEN_TTS_URL}/{voice_id}",
            json=payload,
            headers=headers,
        ) as resp:
            resp.raise_for_status()
            return await resp.read()
```

*The `voice_id` can be any of ElevenLabs’ pre‑built voices or a custom clone you created in the dashboard.* 

Discord expects audio in Opus format inside an `AudioSource`. The easiest way is to pipe the MP3 through FFmpeg and let `discord.py` handle the streaming.

``` php
import discord
import io
import subprocess

def mp3_to_opus(mp3_bytes: bytes) -> discord.FFmpegPCMAudio:
    """
    Takes raw MP3 bytes, writes them to a pipe, and returns an FFmpegPCMAudio source.
    """
    # Use a BytesIO object as a temporary file
    mp3_buffer = io.BytesIO(mp3_bytes)

    # FFmpeg command – reads from stdin (-i pipe:0) and outputs opus to stdout
    ffmpeg_options = {
        "before_options": "-nostdin",
        "options": "-f s16le -ar 48000 -ac 2 pipe:1",
    }

    return discord.FFmpegPCMAudio(mp3_buffer, **ffmpeg_options)
```

Below is a minimal bot that joins a voice channel when you type `!join`, then reads any subsequent message that starts with `!say`.

``` python
import discord
from discord.ext import commands

intents = discord.Intents.default()
intents.message_content = True   # Needed for reading message text

bot = commands.Bot(command_prefix="!", intents=intents)

@bot.event
async def on_ready():
    print(f"🤖 {bot.user} is ready!")

@bot.command()
async def join(ctx):
    """Bot joins the author's voice channel."""
    if ctx.author.voice:
        channel = ctx.author.voice.channel
        await channel.connect()
        await ctx.send(f"Joined {channel.name} 🎤")
    else:
        await ctx.send("You need to be in a voice channel first!")

@bot.command()
async def say(ctx, *, message: str):
    """Bot speaks the supplied text using ElevenLabs."""
    if not ctx.voice_client:
        await ctx.send("I'm not in a voice channel. Use `!join` first.")
        return

    # 1️⃣ Synthesize speech
    try:
        mp3_data = await synthesize(message)
    except Exception as e:
        await ctx.send(f"❗ TTS error: {e}")
        return

    # 2️⃣ Convert to Opus and play
    audio_source = mp3_to_opus(mp3_data)
    ctx.voice_client.play(audio_source, after=lambda e: print("Finished playing", e))

    await ctx.send(f"🗣 Speaking: *{message}*")

@bot.command()
async def leave(ctx):
    """Disconnects the bot from voice."""
    if ctx.voice_client:
        await ctx.voice_client.disconnect()
        await ctx.send("Goodbye! 👋")
    else:
        await ctx.send("I'm not connected to any voice channel.")

bot.run(os.getenv("DISCORD_TOKEN"))
```

`!join``!say <text>``<text>` to ElevenLabs, gets back an MP3, pipes it through FFmpeg, and streams the resulting Opus audio into the channel.
`!leave`
You can expand this skeleton in many ways:

`on_member_join`, `on_message_delete`, etc., and let the bot broadcast them.
`voice_id`, and use it for a truly unique bot personality.
ElevenLabs lets you create a custom voice from as little as 30 seconds of audio. Once you’ve uploaded your sample, you’ll receive a new `voice_id`. Replace the default ID in the `synthesize` function and your bot will speak with *your* voice (or that of a fictional character).

```
CUSTOM_VOICE_ID = "your_custom_voice_id_here"
mp3_data = await synthesize("Welcome to the server!", voice_id=CUSTOM_VOICE_ID)
```

The `voice_settings` payload supports `stability`, `similarity_boost`, `style`, and more. Play around with these values to make the bot sound calm, excited, or even robotic.

```
"voice_settings": {
    "stability": 0.65,
    "similarity_boost": 0.92,
    "style": 0.5,
    "use_speaker_boost": true
}
```

| Symptom | Likely Cause | Fix | 
|---|---|---|
| Bot joins but no audio plays | FFmpeg not in `$PATH` or wrong options | Verify `ffmpeg -version` works and use the`ffmpeg_options` shown above | 
| “TTS error: 401 Unauthorized” | Invalid or missing ElevenLabs API key | Set `ELEVEN_API_KEY` env var correctly | 
| Audio is garbled or too fast | MP3 not being piped correctly | Ensure you’re using `discord.FFmpegPCMAudio` with the proper`before_options` (`-nostdin` ) | 
| Bot disconnects after a few seconds | Rate limit hit on ElevenLabs | Implement a simple in‑memory cooldown (e.g., 1 request per 2 seconds) | 

If you plan to keep the bot running 24/7, consider:

`python:3.11-slim` image with FFmpeg installed (`apt-get update && apt-get install -y ffmpeg`).
`pm2`, `systemd`, or a cloud function that keeps the process alive.
`DISCORD_TOKEN` and `ELEVEN_API_KEY` in environment variables or a secret manager rather than hard‑coding them.
Adding AI voice to a Discord bot is no longer a research‑paper exercise. With a few lines of Python, a free ElevenLabs account, and a dash of creativity, you can turn a silent bot into a conversational companion that reads announcements, narrates games, or simply greets newcomers in a custom‑cloned voice.

Give it a try, experiment with different voice styles, and watch your community react to the new level of immersion.

**Ready to give your bot a voice?** Sign up through this link and start generating high‑quality speech instantly: [https://try.elevenlabs.io/kr07zfuqn1bp](https://try.elevenlabs.io/kr07zfuqn1bp) 

Happy coding, and may your bots always be heard!
