Evaluating S2S Model Quality with Human Preference Data David AI released DAI-S2S-ST, a single-turn speech-to-speech human preference leaderboard built on 153,000 comparative human preference ratings across 7 models, 819 distinct prompts, and 12 questions. The leaderboard asks raters twelve distinct side-by-side questions spanning overall preference, humanness (engagement, emotional appropriateness, naturalness, pleasantness, conversational register), and technical quality, covering eleven categories of consumer-oriented scenarios. David AI said its initial results tell a different story than current S2S leaderboards, and that future work will extend the evaluation to longer, multi-turn interactions. Overview Today we’re releasing DAI-S2S-ST, our single-turn S2S speech-to-speech human preference leaderboard, built on 153k comparative human preference ratings across 7 models, 819 distinct prompts, and 12 questions. For this evaluation, we grade model responses to single-turn, pre-recorded input prompts. Our initial results tell a different story than current S2S leaderboards. We believe the future of voice AI is one where people spend hours a day interacting with models across apps, devices, and new physical interfaces. For that to happen, models need to be more than capable. They need to be compelling enough that people actually want to keep talking to them. That is different from how voice assistants have traditionally been used. Most were built for short transactions: set a timer, check the weather, play a song. If the model understands the request and completes the task then the interaction is successful, even if the voice sounds robotic or awkward. Interacting with models for hours a day raises the bar — and requires us to evaluate models differently. Naturalness, personality, empathy, and the overall quality of the interaction become critical. Most voice benchmarks measure whether the model understood the input, answered correctly, or completed the task. These questions matter, but they don’t tell us whether someone would actually want to keep talking to the model. To measure that, we need to ask humans. DAI-S2S-ST is a first step toward measuring not just what a model can do, but what it feels like to interact with one: 1. Human preference across 12 dimensions. Existing human preference leaderboards e.g., Voice Showdown https://web.archive.org/web/20260429190552/https://labs.scale.com/blog/voice-showdown , now retired, and Speech Agent Arena https://artificialanalysis.ai/speech-to-speech/arena typically collapse preference into a single overall judgment. DAI-S2S-ST asks twelve distinct side-by-side questions to understand what actually drives preference. 2. Focused on consumer use cases. We evaluate eleven categories of consumer-oriented scenarios, where the qualities that drive a good interaction differ from what most task-oriented voice-agent leaderboards cover e.g., τ³-Voice https://taubench.com/leaderboard/?benchmark=voice and Eva Bench https://servicenow.github.io/eva/ results . This initial version focuses on single-turn interactions. Future work will extend this evaluation to longer, multi-turn interactions. What’s In the Evaluation For full details on our methodology, see the methodology section methodology . Our US-based voice contributor network recorded in a variety of environments that simulate real user conversations. 2. Evaluation Each recording was run through inference with each of the relevant models, and then David AI’s paid rater panel completed side-by-side preference judgements against our rubric. These questions are slightly modified from those in our LALM-as-judge vs HITL https://research.withdavid.ai/blog/lalm-as-judge-vs-hitl post based on rater feedback and analysis of those earlier results. Additional details on the evaluation methodology for this leaderboard can be found in the methodology section. | Rubric questions | | | |---|---|---| | Axis | Parameter | Rater question | | Overall | Overall | “Overall, which response do you prefer?” | | Humanness | Engagement | “Which response was more engaging?” | | | Emotional appropriateness