Today we’re introducing Audio Realism Bench, our benchmark for AI conversational realism.
Modern TTS has saturated the traditional 1–5 mean opinion score (MOS) and while it is able to capture intelligibility and naturalness, it no longer meaningfully separates frontier systems that all produce clean speech.
Audio Realism Bench instead evaluates realism as an emerging signal for high quality audio generations. Given two recordings of the same transcript, (presented blindly and in randomized order) our team of vetted native speakers choose which voice sounds more like a real person. We include real human recordings in the same pool as well, so beyond the headline Elo we can measure how often each model is picked over a real person.
Rather than measuring synthesis quality alone, Audio Realism Bench evaluates conversational realism across the kinds of speech users increasingly deploy in production systems, including phone agents, conversational assistants, and spoken explainers.
Why Realism Must be Measured
The generation of artificial speech has shifted from content generation toward human-facing interaction. Modern voice systems increasingly answer customer support calls, conduct scheduling conversations, guide users through products, and participate in real-time dialogue.
Existing audio benchmarks, including our own AudioAgentBench, primarily evaluate quantitative attributes such as pronunciation accuracy, prompt adherence, latency, and speech quality. While these remain important measurements, none test the speech quality qualitatively and whether or not it sounds like a person.
Especially when deployed in human facing environments, generation effectiveness depends on whether speech exhibits the subtle characteristics people associate with human conversation: timing, hesitation, breathing, disfluencies, conversational pacing, and other cues that listeners evaluate subconsciously.
By extending the audio evaluation space to account for realism, corporations are able to understand where voice agents are likely to succeed, where human oversight remains appropriate, and where authentication systems that rely primarily on voice may become increasingly vulnerable. Likewise, identifying the gaps that exist in model speech pattern helps to identify capability gaps and where human interference is still required.
How We Measure Realism #
The fundamental unit of scoring is the pairwise comparison. A listener is presented with two clips of the same prompt from two models, blind to their identity and with left/right order randomized, and selects the one that sounds more human. Each selection constitutes a single vote.
These votes are aggregated with the Bradley-Terry model, the maximum-likelihood method behind Elo-style ratings, which resolves every head-to-head win and loss into a single Elo rating per model. A higher rating means the model is chosen as more human more often.
For each prompt we also record a human reading it and enter that clip into the pool as a human baseline. Every model is then measured directly against that baseline: how often it is picked over a real person.
What We Test On #
We score models from a fixed, versioned set of prompts, most up to roughly 40 words (about 15 seconds), though phone-agent prompts are single turns and tend to be shorter. They are organized into three conversation categories defined by speech register (how someone is talking) rather than topic; within each category we deliberately vary the subject matter. Because a category of conversation follows how something is spoken rather than what it is about, one category can span more than one topic: a healthcare call to book an appointment is a phone agent, while a walkthrough of what a medication does is an explainer.
Within each conversation category, prompts are stratified across speech difficulty, so that each model is evaluated on both smooth, predictable speech and harder, less predictable passages that are more difficult to voice naturally.
Acting-heavy speech (game characters, dramatized fiction, persona-driven companion voices) is out of scope for this version, deferred to a future expressive track. Transcripts keep the hesitations, false starts, and fillers of real speech wherever the register allows. Numbers, times, and amounts are written as spoken words rather than digits or symbols, so the transcripts read exactly as a model is expected to voice them.
Our pool of samples is comprised of 500+ prompts, kept fully private and never published. We are aware there may be preexisting bias from certain age groups or backgrounds who may have been more exposed to certain model voices. Therefore, we use experts across a range of age groups, backgrounds, and experience with AI in order to debias this effect.
Conversation Categories #
Phone Agents
Handling real-time spoken customer support, inquiries, and routing.
Sample Text: “Sure, I can help you with that, um, what is the payment amount?”
Bland Speech v3
Explainers
Explaining, teaching, or clarifying an idea, process, or event by a single-speaker
Sample Text: “Let me, you know, some of us, have a guide, the library staff, time and time again to find that book, and some of them do it with greater detail than others, but let me just say that this not a challenging task, we are certain.”
Eleven Labs
Conversations
Podcasts and interviews, the most casual, unscripted register
Sample Text: “That's amazing, that's amazing because they want to, you know, you know, build the park, plant trees on a hill, all of your land, and then to do that, that's not what we want. We're going to develop this area sustainably.”
Gemini 3.1 Flash
One Voice Per Model, Per Gender #
Most providers ship many voices, so the voice can matter as much as the model. For each model we use one female and one male voice, chosen as its most naturally conversational American-English voice. Candidates must be the right gender and American-accented; among those we take the voice the provider’s docs describe as most conversational — preferring, in order, conversational/voice-agent → warm/natural → casual → neutral.
For gender, each matchup draws a random prompt and a random gender, and both models (and the human anchor) speak it with their selected voice of that gender. Every comparison is same-gender, so gender is never the difference between the two clips. As expected, human voices lead our rankings.
Second only to human speech is Bland Speech v3 by Bland AI, with an Elo of 1365. MAI-Voice-2 by Microsoft AI takes the next spot at 1224 Elo, followed by Grok TTS by SpaceXAI at 1156 Elo.
Bland Speech v3 is the latest model by Bland AI, which they claim “to be the world’s first Human Speech Engine (HSE)”. Trained on over 20m+ real conversations, Bland Speech v3 makes the first step at generating vocal assets that don’t just sound clean, but sound realistically human.
Our evaluation is an open, evolving process and we genuinely welcome feedback, suggestions, and new models to consider.
To propose a model, provider, or new subcategory for a future run, reach us at [email protected].