# Evaluating S2S Model Quality with Human Preference Data

> Source: <https://research.withdavid.ai/blog/st-s2s-eval>
> Published: 2026-09-17 01:42:28+00:00

## Overview

Today we’re releasing DAI-S2S-ST, our single-turn S2S (speech-to-speech) human preference leaderboard, built on 153k comparative human preference ratings across 7 models, 819 distinct prompts, and 12 questions. For this evaluation, we grade model responses to single-turn, pre-recorded input prompts. Our initial results tell a different story than current S2S leaderboards.

We believe the future of voice AI is one where people spend hours a day interacting with models across apps, devices, and new physical interfaces. For that to happen, models need to be more than capable. They need to be compelling enough that people actually want to keep talking to them.

That is different from how voice assistants have traditionally been used. Most were built for short transactions: set a timer, check the weather, play a song. If the model understands the request and completes the task then the interaction is successful, even if the voice sounds robotic or awkward.

Interacting with models for hours a day raises the bar — and requires us to evaluate models differently. Naturalness, personality, empathy, and the overall quality of the interaction become critical. Most voice benchmarks measure whether the model understood the input, answered correctly, or completed the task. These questions matter, but they don’t tell us whether someone would actually want to keep talking to the model. To measure that, we need to ask humans.

DAI-S2S-ST is a first step toward measuring not just what a model can do, but what it feels like to interact with one:

1. **Human preference across 12 dimensions.** Existing human preference leaderboards (e.g.,[Voice Showdown](https://web.archive.org/web/20260429190552/https://labs.scale.com/blog/voice-showdown) , now retired, and[Speech Agent Arena](https://artificialanalysis.ai/speech-to-speech/arena) ) typically collapse preference into a single overall judgment. DAI-S2S-ST asks twelve distinct side-by-side questions to understand what actually drives preference.
2. **Focused on consumer use cases.** We evaluate eleven categories of consumer-oriented scenarios, where the qualities that drive a good interaction differ from what most task-oriented voice-agent leaderboards cover (e.g.,[τ³-Voice](https://taubench.com/leaderboard/?benchmark=voice) and[Eva Bench](https://servicenow.github.io/eva/#results) ).

This initial version focuses on single-turn interactions. Future work will extend this evaluation to longer, multi-turn interactions.

## What’s In the Evaluation

For full details on our methodology, see [the methodology section](#methodology).

Our US-based voice contributor network recorded in a variety of environments that simulate real user conversations.

### 2. Evaluation

Each recording was run through inference with each of the relevant models, and then David AI’s paid rater panel completed side-by-side preference judgements against our rubric.

These questions are slightly modified from those in our [LALM-as-judge vs HITL](https://research.withdavid.ai/blog/lalm-as-judge-vs-hitl) post based on rater feedback and analysis of those earlier results. Additional details on the evaluation methodology for this leaderboard can be found in the methodology section.

| Rubric questions |  |  | 
|---|---|---|
| Axis | Parameter | Rater question | 
| Overall | Overall | “Overall, which response do you prefer?” | 
| Humanness | Engagement | “Which response was more engaging?” | 
|  | Emotional appropriateness <sup>[*](#sts-gated-rubrics-note)</sup> | “Which response’s emotional tone responded more appropriately to the user’s emotional state and the situation?” | 
|  | Naturalness | “Which response’s delivery sounded more like a natural speaker?” | 
|  | Pleasantness | “Which response was more pleasant to listen to?” | 
|  | Conversational register | “Which response used more natural, spoken-conversation language, rather than sounding like written text being read aloud?” | 
| Technical Quality | Response quality | “Which response was more helpful in achieving what the user was trying to accomplish?” | 
|  | Reasoning <sup>[*](#sts-gated-rubrics-note)</sup> | “Which response’s reasoning — when explaining its thinking or answering a question requiring judgment — was more correct and logically sound?” | 
|  | Instruction following <sup>[*](#sts-gated-rubrics-note)</sup> | “Which response followed the specific instructions or constraints more closely?” | 
|  | Perceived audio quality | “Which response had better audio quality?” | 
|  | Pronunciation accuracy | “Which response had more accurate pronunciation?” | 
|  | Voice consistency | “Which response kept a more consistent voice throughout (less noticeable change in character, timbre, or quality over time)?” | 

## Results

### Current leaderboards do not capture naturalness or acoustic quality

Because of our focus on (1) human perception of conversation quality and (2) consumer-oriented scenarios, our overall rankings diverge from those produced by other public leaderboards with coverage over the same models.

For example, Grok Think Fast 2.0 and Qwen3 both perform very well on [leaderboards that measure task completion](https://taubench.com/leaderboard/?benchmark=voice) on transactional voice-agent scenarios with verifiable outcomes, but perform significantly worse on our leaderboard.

Our rubrics resolve into two distinct groups, which we have labeled “Humanness” and “Technical Quality”. We found them by running hierarchical clustering over per-rubric rater responses (after adjusting for rater- and prompt-level effects — see [methodology section](#methodology)), which gave a stable clustering at k=2 that consistently reproduced across rater-prompt bootstrapping:

- **Humanness** : conversational register, emotional appropriateness, engagement, naturalness, pleasantness
- **Technical Quality** : instruction following, perceived audio quality, pronunciation accuracy, reasoning, response quality, voice consistency

In [previous experiments](https://research.withdavid.ai/blog/lalm-as-judge-vs-hitl) we grouped questions into “content”, “acoustic” and “interaction” categories. However, during this analysis, we found that these groupings did not capture meaningfully different information — not only did they produce the same model stack ranks, they produced Elo distributions that were statistically indistinguishable from the “overall” human preference Elo.

#### CMOS Elo by rubric metric

| CMOS Elo per model per rubric metric, with 95% confidence intervals |  |  |  |  |  |  |  | 
|---|---|---|---|---|---|---|---|
| Metric | GPT-Realtime 2.1 | GPT-Realtime 2.0 | Gemini 3.1 Flash Live | Grok voice-think-fast 2.0 | Nova 2 Sonic | Cascade | Qwen3-Omni | 
|---|---|---|---|---|---|---|---|
| Conversational register | 1160 [1140, 1179] | 1131 [1109, 1152] | 1122 [1100, 1143] | 981 [961, 1002] | 921 [888, 956] | 844 [808, 875] | 841 [798, 880] | 
| Engagement | 1138 [1118, 1155] | 1102 [1080, 1120] | 1105 [1080, 1133] | 971 [952, 990] | 988 [957, 1029] | 883 [853, 912] | 813 [782, 847] | 
| Emotional appropriateness | 1179 [1157, 1208] | 1138 [1104, 1173] | 1073 [1044, 1105] | 977 [950, 1008] | 964 [921, 1003] | 866 [811, 916] | 802 [747, 847] | 
| Naturalness | 1176 [1160, 1196] | 1149 [1128, 1170] | 1127 [1106, 1147] | 977 [954, 998] | 960 [935, 988] | 810 [771, 844] | 800 [759, 840] | 
| Pleasantness | 1164 [1146, 1183] | 1134 [1115, 1155] | 1107 [1085, 1132] | 987 [966, 1005] | 978 [950, 1003] | 847 [809, 885] | 783 [748, 821] | 
| Response quality | 1113 [1099, 1128] | 1105 [1087, 1125] | 1009 [991, 1027] | 1006 [990, 1021] | 1050 [1020, 1080] | 947 [918, 974] | 770 [742, 798] | 
| Instruction following | 1075 [1038, 1106] | 1064 [1023, 1103] | 1015 [982, 1049] | 1015 [979, 1048] | 966 [902, 1026] | 1003 [940, 1083] | 861 [815, 914] | 
| Reasoning | 1118 [1095, 1145] | 1103 [1073, 1137] | 979 [949, 1008] | 1007 [978, 1037] | 1109 [1055, 1162] | 927 [870, 981] | 757 [700, 809] | 
| Perceived audio quality | 1071 [1048, 1093] | 1068 [1047, 1090] | 1103 [1082, 1126] | 1015 [993, 1039] | 915 [883, 949] | 990 [947, 1033] | 837 [796, 873] | 
| Pronunciation accuracy | 1056 [1035, 1079] | 1064 [1038, 1087] | 1073 [1053, 1094] | 1018 [996, 1040] | 1026 [1000, 1048] | 1002 [964, 1039] | 762 [716, 802] | 
| Voice consistency | 1040 [1015, 1062] | 1066 [1041, 1087] | 1039 [1013, 1068] | 1023 [999, 1044] | 1020 [992, 1048] | 985 [949, 1022] | 827 [781, 875] | 

GPT-Realtime 2.1Overall CMOS Elo: 1149

The root cause of the correlation of these dimensions across models is likely traceable to the model training pipeline: perhaps encoded in the inductive bias of modern architectures, influenced by related latent patterns in the training data, or implicit in the internal metrics different research labs are hill climbing.

These two clusters are a statement about the variation in performance of current models on our prompt distribution. As models improve unevenly across capabilities, or as the use cases evaluated change, these relationships may separate. It is entirely possible that a future model will excel in perceived audio quality and voice consistency but struggle at instruction following and reasoning, just as it is entirely possible that transactional use cases show no strong relationship between human preference and the “humanness” of the model.

### Raters prefer certain voices over others

Further evidence for the importance of aesthetics to human preference comes from examining model performance disaggregated by speaking voice. By default, each model is evaluated using two of its supported speaking voices (one male and one female).

Sometimes, as was the case with Grok Ara vs. Rex voices, this preference was very prominent.

Users display a strong preference for Ara on Humanness dimensions, and a weaker preference for Ara on Technical Quality dimensions.

#### Ara’s margin over Rex (Grok voice-think-fast 2.0)

| Grok Ara against Rex by dimension: ratings, CMOS margin, 95% confidence interval and win rate. |  |  |  |  |  | 
|---|---|---|---|---|---|
| Dimension | Ratings | Ara’s margin |  | Ara win rate |  | 
|---|---|---|---|---|---|
|  |  | CMOS points | 95% CI | ties excluded |  | 
| Overall |  |  |  |  |  | 
| Overall preference | 446 | +0.93 | [0.63, 1.21] | 74% |  | 
| Humanness |  |  |  |  |  | 
| Pleasantness | 446 | +1.03 | [0.77, 1.29] | 82% |  | 
| Naturalness | 446 | +0.98 | [0.74, 1.22] | 83% |  | 
| Conversational register | 446 | +0.81 | [0.58, 1.05] | 79% |  | 
| Engagement | 446 | +0.76 | [0.52, 1.00] | 78% |  | 
| Emotional appropriateness | 126 | +0.72 | [0.28, 1.14] | 74% |  | 
| Technical Quality |  |  |  |  |  | 
| Instruction following | 78 | +0.40 | [-0.05, 0.88]n.s. | 78% |  | 
| Reasoning | 122 | +0.36 | [-0.08, 0.78]n.s. | 63% |  | 
| Response quality | 446 | +0.32 | [0.08, 0.56] | 64% |  | 
| Perceived audio quality | 446 | +0.19 | [-0.03, 0.40]n.s. | 62% |  | 
| Pronunciation accuracy | 446 | +0.18 | [0.02, 0.35] | 62% |  | 
| Voice consistency | 446 | +0.14 | [-0.08, 0.37]n.s. | 55% |  | 

Notably, three Technical Quality dimensions should not vary by voice: reasoning, instruction following, and response quality. In most modern S2S architectures, one model backs every voice, so the underlying quality should be the same. For two models, our data disagrees; Grok and GPT-Realtime both show a voice gap on these dimensions. Raters prefer Grok’s Ara voice on response quality, and its reasoning and instruction-following margins lean the same way. GPT-Realtime shows the same lean between its voices. The reasoning and instruction-following results should be interpreted cautiously; they are gated (applying to only some prompts) and have smaller samples with wide confidence intervals. It’s possible this dynamic is caused by a halo effect from the Humanness preference, or that some models have deeper architectural or training data distribution explanations for the differences in voice performance.

Regardless of the underlying reason, these results demonstrate that you cannot assume comparable performance across model voices as most S2S leaderboards today do. Moreover, they demonstrate that voice-specific factors can strongly influence overall human preference on model outputs.

### Higher thinking level does not increase performance on our leaderboard

In contrast to speaking voices, our data showed no material difference in overall model performance based on model thinking level:

This pattern holds even in isolated dimensions where additional thinking would be most likely to differentiate performance. Higher-thinking variants did not meaningfully out-perform their lower-thinking counterparts in instruction following or reasoning, for example.

We attribute the lack of differentiation among thinking modes primarily to the nature of our task distribution: the single-turn mix does not contain the sort of difficult multi-step problems where you would expect extra thinking to pay off. However, it is notable that the overall differentiation of model thinking levels across S2S leaderboards is mixed — Speech Agent Arena shows relatively little difference between different thinking levels, while τ³-Voice shows a larger effect, but only for certain models. Overall, the public evidence suggests that higher thinking can improve specific reasoning capabilities without reliably improving broader S2S model performance.

## Conclusion

Our findings demonstrate that human preference is shaped as much by communication style as by content. This analysis relies on preference data from 283 participants, underscoring the value of human evaluation for measuring outcomes that are difficult to quantify.

At David AI we continue to evaluate public models, pre-release checkpoints, and agentic speech-to-speech workflows using a mixture of human and automated evaluation to better understand real-world interaction quality.

For inquiries regarding model evaluation or future research collaborations, contact [evals@withdavid.ai](mailto:evals@withdavid.ai). We are also currently [hiring](https://www.withdavid.ai/jobs) for the Evaluations team.

DAI-S2S-ST: Comparative Mean Opinion Score (CMOS)

| CMOS Elo for every model on every rubric metric. Selecting a metric ranks it in the board below. |  |  |  |  |  |  |  |  |  |  |  |  | 
|---|---|---|---|---|---|---|---|---|---|---|---|---|
|  | Overall | Humanness |  |  |  |  | Technical Quality |  |  |  |  |  | 
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model |  |  |  |  |  |  |  |  |  |  |  |  | 
| GPT-Realtime 2.1 |  |  |  |  |  |  |  |  |  |  |  |  | 
| GPT-Realtime 2.0 |  |  |  |  |  |  |  |  |  |  |  |  | 
| Gemini 3.1 Flash Live |  |  |  |  |  |  |  |  |  |  |  |  | 
| Grok voice-think-fast 2.0 |  |  |  |  |  |  |  |  |  |  |  |  | 
| Nova 2 Sonic |  |  |  |  |  |  |  |  |  |  |  |  | 
| Cascade |  |  |  |  |  |  |  |  |  |  |  |  | 
| Qwen3-Omni |  |  |  |  |  |  |  |  |  |  |  |  | 

Per category breakdown

“Overall, which response do you prefer?”

Showing Overall preference, ranked by CMOS Elo. Higher is better.

| Overall preference rankings by CMOS Elo across 7 models, with 95% confidence intervals. |  |  |  |  | 
|---|---|---|---|---|
| # |  |  | 95% CI | Interval plot | 
|---|---|---|---|---|
| 1 | GPT-Realtime 2.1 | 1149 | [1130, 1168] |  | 
| 2 | GPT-Realtime 2.0 | 1128 | [1109, 1149] |  | 
| 3 | Gemini 3.1 Flash Live | 1071 | [1050, 1091] |  | 
| 4 | Grok voice-think-fast 2.0 | 993 | [977, 1009] |  | 
| 5 | Nova 2 Sonic | 979 | [948, 1014] |  | 
| 6 | Cascade | 897 | [864, 928] |  | 
| 7 | Qwen3-Omni | 784 | [751, 817] |  | 

## Methodology

### Overview

The David AI S2S Single Turn (DAI-S2S-ST) leaderboard is a benchmark that evaluates S2S model quality on single-turn human-recorded audio prompts focused on real consumer use cases. Quality is evaluated by human comparative preference ratings across 12 dimensions of quality (11 specific dimensions + 1 “overall” rating), which are grouped into “Humanness” and “Technical Quality” categories for analysis. Comparative ratings are used to fit a Bradley-Terry Elo model, and confidence intervals are reported for each Elo score.

Two voices and two thinking levels (where supported) were included for each evaluated model (except the cascade baseline). These voices and thinking levels are aggregated into a single model-level Elo on the leaderboard; the voice and thinking-level comparisons are reported as their own figures in the results section. The audio prompts themselves are kept private to mitigate eval set leakage risk - the audio samples included in the blog post are representative, but not used in the leaderboard itself.

### Audio prompt generation

Audio prompt scripts are drawn from eleven Objective Categories and carry a difficulty level from L1 (easiest) to L5 (hardest), plus a separate prompt set captured under degraded acoustic conditions. Each script is then recorded by a human contributor on the David AI audio data collection platform. The samples used in the leaderboard consist of ~14 input prompt scripts for each difficulty level x objective category pair for a total of 819 audio prompts (11 objective categories x 5 difficulty levels x ~14 prompts + ~45 degraded audio samples).

| Table 1:  Audio prompt count by difficulty level, with the degraded-condition set counted separately. |  | 
|---|---|
| Difficulty | Prompts | 
|---|---|
| L1 | 156 | 
| L2 | 157 | 
| L3 | 160 | 
| L4 | 156 | 
| L5 | 145 | 
| L1–L5 core | 774 | 
| Degraded | 45 | 
| Total | 819 | 

**Table 1:** Audio prompt count by difficulty level, with the degraded-condition set counted separately.

The objective categories and audio samples are described in the main blog post.

Each script was sent to David AI voice contributors to be recorded on our audio platform. Each script was recorded by at least one contributor in the target acoustic environment: clean, meaning minimal background noise and a near-field microphone setup; or degraded, meaning indoor background noise, outdoor background noise, or a far-field microphone setup. The recorded samples were passed through an automated and human review QA process, and one audio sample was selected for each unique script. The selected samples were volume normalized and trimmed, but otherwise unprocessed.

### Model selection

We surveyed the available S2S leaderboards in July 2026 and identified what we believed to be the strongest S2S model offered by each of the SOTA S2S model providers available at the time. For model families that supported multiple thinking modes, we selected the highest and lowest available variants for comparison.  In practice, our analysis shows negligible performance differences on our eval set for different thinking levels, and our results are aggregated across thinking levels in our leaderboard for simplicity (the comparison by thinking level is reported in the [results section](#results)).

In addition, we identified a SOTA cascade model pipeline and a high-performing open-source model to serve as comparison baselines. The cascade pipeline uses Scribe-V2 for speech-to-text (STT) transcription, GPT-5.6 Terra as base language model, and ElevenLabs Flash v2.5 for text-to-speech (TTS). These models were chosen as a balance of quality and latency, as the highest audio quality TTS models have high latency and are not a realistic comparison point for deployable S2S systems.

Finally, we included GPT-Realtime-2 (xhigh thinking) in addition to the GPT-Realtime-2.1 variants because GPT-Realtime-2.1 was [framed by the developers](https://community.openai.com/t/new-realtime-models-on-the-api-gpt-realtime-2-1-and-gpt-realtime-2-1-mini/1385896) as an incremental update with specific feature support and latency optimization; we considered it plausible that GPT-Realtime-2 might outperform 2.1 on human preference over our eval set once latency differences were removed. In practice we see only a minimal performance difference across GPT-Realtime-2 and GPT-Realtime-2.1.

The final list of models evaluated at launch time includes:

| Table 2:  Models evaluated at launch, with the thinking level each was run at. |  | 
|---|---|
| Model | Thinking | 
|---|---|
| Gemini 3.1 Flash Live | minimal | 
|  | high | 
| GPT-Realtime-2.1 | minimal | 
|  | xhigh | 
| GPT-Realtime-2 | xhigh | 
| Grok voice-think-fast 2.0 | high | 
|  | none | 
| Nova 2 Sonic | default | 
| Cascade | default (GPT-5.6-Terra) | 
| Qwen3-Omni-30B-A3B-Instruct | default | 

**Table 2:** Models evaluated at launch, with the thinking level each was run at.

We expect to add more models and variants to this leaderboard over time. The full list of 19 model configurations is available in the Model Configuration Details section.

### Voice selection

We selected two voices (one male, one female) for each S2S model except the cascade baseline. The data for each voice is aggregated in our leaderboard for simplicity; the comparison between a model’s two voices is reported in the results section.

To minimize the impact of voice preference on the ranking, we prioritized selecting voices that are:

- Neutral EN-US accented
- Medium pitch range within each gender (eliminating any voice with a noticeably high or low pitch)
- No strong sociolinguistic markers (eliminating any voice that is noticeably or characteristically from a particular sociological group)

When a model had a default voice identified in the documentation that fit the above criteria, we used that voice. When the default voice did not fit the criteria, or when there was no default voice specified, we selected a random voice that met the criteria.

The model that was most impacted by this methodology was Grok voice-think-fast 2.0, which uses a default voice with an EN-GB accent (Eve) that we replaced with another voice (Ara).

The final voices used for each model as of launch time were:

| Table 3:  The selected voice for each model. |  |  | 
|---|---|---|
| Model | Voice | Gender | 
|---|---|---|
| Gemini 3.1 Flash Live | Zephyr | F | 
|  | Iapetus | M | 
| GPT-Realtime-2.1 / GPT-Realtime-2 | Marin | F | 
|  | Cedar | M | 
| Grok voice-think-fast 2.0 | Ara | F | 
|  | Rex | M | 
| Nova 2 Sonic | Tiffany | F | 
|  | Matthew | M | 
| Cascade | Rachel | F | 
| Qwen3-Omni-30B-A3B-Instruct | Ethan | M | 
|  | Chelsie | F | 

**Table 3:** The selected voice for each model.

### Model Inference

Each combination of <model, thinking level, voice> was prompted with each of the 819 audio input prompts using their officially recommended inference settings to the extent possible.

We did not hand-tune system prompts, user instruction, or any other parameter settings for individual models.  For models that provide a default prompt or other recommended settings in their public documentation, we used these default and recommended values - see the Qwen case below.  For models that did not provide a default prompt or recommended settings, we used a standard instruction of “`You are a helpful voice assistant`” and used the default parameter values.  The full set of model inference configuration values can be found in the Model Configuration Details section.

The model configuration that stands out as the most unusual is the system prompt used for the Qwen3 Omni model:

```
You are a virtual voice assistant with no gender or age. You are communicating with the user. In user messages, "I/me/my/we/our" refer to the user and "you/your" refer to the assistant. In your replies, address the user as "you/your" and yourself as "I/me/my"; never mirror the user's pronouns—always shift perspective. Keep original pronouns only in direct quotes; if a reference is unclear, ask a brief clarifying question. Interact with users using short (no more than 50 words), brief, straightforward language, maintaining a natural tone. Never use formal phrasing, mechanical expressions, bullet points, overly structured language. Your output must consist only of the spoken content you want the user to hear. Do not include any descriptions of actions, emotions, sounds, or voice changes. Do not use asterisks, brackets, parentheses, or any other symbols to indicate tone or actions. You must answer users' audio or text questions, do not directly describe the video content. You should communicate in the same language strictly as the user unless they request otherwise. When you are uncertain (e.g., you can't see/hear clearly, don't understand, or the user makes a comment rather than asking a question), use appropriate questions to guide the user to continue the conversation. Keep replies concise and conversational, as if talking face-to-face
```

We initially implemented Qwen3 with the same “`You are a helpful voice assistant`” system prompt used for most other models, but we saw a high rate of catastrophic failure where the model would respond with incoherent content.  After investigation, the [Qwen3 Omni GitHub page](https://github.com/QwenLM/Qwen3-Omni) includes a [recommended prompt](https://github.com/QwenLM/Qwen3-Omni#prompt-for-audio-visual-interaction) for “Audio-Visual interaction... where the input consists of a video and its corresponding audio.” While video is not provided in our audio input prompts, we found model responses much more stable using this prompt, and used this to generate the actual evaluation samples.

Audio requests were streamed to standard model inference API endpoints, and the audio responses were collected after the end of turn event was emitted by the model API. For models that provide parallel output streams to the audio, everything but the final audio output was ignored for purposes of rating. Specifically:

- Some models provide a “provisional” audio stream that outputs audio chunks as they become available while the final audio stream is still being constructed.
- Some models provide a separate “thinking” token stream.
- Some models provide an input and/or output transcript in addition to the output audio stream.

All of these values were recorded for potential future analysis, but ignored for the purposes of human evaluation.

### Human Preference Rating Collection

A rating pair consists of the output from one specific model configuration <model, thinking level, voice> paired with the output from another configuration on the same input prompt. Rating pairs can include output from the same base model at different thinking levels and/or using different voices. A given model configuration is not compared against all other model configurations for every single input prompt; however, every model configuration had at least partial coverage over other configurations for every input prompt, and every model configuration was paired with every other model configuration at roughly equal rates across all rating pairs. At launch, 19 unique model configurations were evaluated.

For each rating pair, one or more raters were assigned to answer a set of up to 12 distinct questions (or “rubrics”), one of which was “overall preference”. The median number of raters for a given rating pair was 2 (i.e., the typical rating pair was rated by 2 different humans). 3 of the 12 rubrics are “gated,” meaning the question is only asked when it is applicable to the specific input prompt (for example, the “reasoning” question was only asked for input prompts tagged as requiring some amount of reasoning in the response). The full list of rubric questions is detailed in the [main blog post](#rubric-questions).

For CMOS ratings, a single rater was asked to answer all relevant rubric questions for a given rating pair in succession. During collection, questions were grouped into three categories, and each category was answered in full before moving on to the next category:

| Table 4:  The three groups the rubric questions were asked in during collection. |  | 
|---|---|
| Category | Rubric | 
|---|---|
| Content Quality | response quality, instruction following, reasoning, conversational register | 
| Acoustic Quality | audio quality, naturalness, pleasantness, pronunciation, voice consistency | 
| Interaction Quality | emotional appropriateness, engagement | 

**Table 4:** The three groups the rubric questions were asked in during collection.

These category groupings were reevaluated in the analysis stage, and statistical clustering regrouped them into the Humanness and Technical Quality categories as discussed in the main blog post.

MOS ratings were collected in addition to CMOS ratings for internal analysis, but the leaderboard reports only CMOS results.

### Model Clustering Methodology

To group the rubric dimensions into categories, we cluster rubrics by the similarity of their per-model performance profiles. For each rubric, we take the per-model latent strengths from the same ordinal Bradley-Terry fit used for the leaderboard, yielding a rubric × model matrix. Two corrections are applied before clustering:

1. Because every rubric's scores are dominated by a shared "general quality" axis (the halo effect), naive clustering would group rubrics by how strongly they track overall quality rather than by what they measure; we therefore project out the first principal component of the (row-centered) matrix and cluster on each rubric's residual profile.
2. Because rubric pairs are scored by the same raters on the same prompts, their observed similarity includes a shared-noise component; we estimate this via a rater-and-prompt clustered bootstrap (1,000 resamples) and subtract it from each pairwise similarity. We then apply average-linkage hierarchical clustering to the corrected similarities.

We assess cluster reliability as the fraction of bootstrap resamples that reproduce the identical partition, against a chance baseline computed on signal-removed data. The two-cluster split of Humanness / Technical Quality reproduces in 82% of resamples (chance rate < 1%); a three-cluster split, which would separate the acoustic rubrics into their own group, reproduces in only 38% and is not reported.

Clustering was computed with GPT-Realtime-2 and 2.1 pooled as a single family (six model families); given their near-identical rubric profiles we do not expect the split to alter the structure.

As a held-out check, we also clustered two disjoint random halves of the rater pool independently: the two-cluster structure and the Humanness core are consistently recovered (mean Adjusted Rand Index 0.60 over 500 splits), though boundary rubrics (instruction following, perceived audio quality) occasionally migrate between clusters. We therefore report the k=2 grouping and treat finer sub-structure as provisional pending more data.

### Model Ranking Methodology

The leaderboard uses CMOS preference scores to directly compute Elo from a Bradley-Terry model. For the primary leaderboard, we use the CMOS scores from the Overall question per model; for fine-grained analysis, we compute individual rankings across rubric categories, as well as per-voice and per-thinking-level variants.

Since CMOS score is an ordinal scale from [-3, 3], we use ordinal Bradley-Terry as the primary model to compute the latent score for each candidate. We also fit two other ranking engines: Massey ranking and classic Bradley-Terry with partial-margin extension (-3..+3 CMOS values normalized to -1..+1 and treated as partial win margin) for cross-method validation. All three engines produce near-identical rankings with each other (Spearman correlation at or above 0.99).

The latent Elo score for each candidate is anchored at 1000 for ease of reading, with a 400-point gap corresponding to 10-to-1 odds:

`Elo = 1000 + (400 / ln 10) · s`

Confidence intervals are calculated using rater-prompt-clustered bootstrapping: we randomly sample raters and prompts with replacement to build a subset and calculate ranking statistics. We repeat the process 1,000 times and report the 95% CI as the range between the 2.5th and 97.5th percentiles of the spread.

## Appendix

## **Extended Data and Robustness Analysis**

### Evaluation Data

At launch, our leaderboard collection consists of ~153k comparative preference ratings from a pool of 283 raters across model responses to 819 human-recorded audio prompts. A single comparative preference rating is defined as a standard CMOS (-3 to +3 with ties allowed) rating for one of 12 rubric questions for a given pair of audio samples. A rated audio sample pair consists of the output of two model configurations that have been fed the same single-turn human-recorded audio input prompts. A model configuration is defined as the combination of <model, voice, thinking level>, and 19 distinct model configurations were evaluated (generally 2 voices and 2 thinking levels for each model, with a few exceptions in the model selection section). Not all model configurations were compared against all other model configurations for a given input prompt, but every model configuration had at least one comparison against another model configuration for each input prompt.

In addition to CMOS, we also collected MOS ratings for internal validation and comparative analysis with CMOS data. This MOS data is not included in the leaderboard or methodology section, but is discussed throughout the appendix.

The rater panel that supplied these ratings was roughly balanced on gender, with a mean age of 40.9 and coverage across age bands from 18 to 65+. Per collection, 283 raters took part in the CMOS collection and 251 raters in the MOS collection; 166 raters did both MOS and CMOS evaluations.

| Table A1:  Rater pool gender and age distribution |  |  |  |  |  |  |  |  | 
|---|---|---|---|---|---|---|---|---|
| Gender | Pool | 18-25 | 26-35 | 36-45 | 46-55 | 56-65 | 65+ | Total | 
|---|---|---|---|---|---|---|---|---|
| M | MOS | 23 | 33 | 28 | 19 | 14 | 4 | 121 | 
|  | CMOS | 25 | 42 | 33 | 28 | 14 | 3 | 145 | 
| F | MOS | 12 | 25 | 31 | 27 | 29 | 2 | 126 | 
|  | CMOS | 12 | 27 | 32 | 29 | 29 | 4 | 133 | 
| Other | MOS | 1 | 3 | 0 | 0 | 0 | 0 | 4 | 
|  | CMOS | 1 | 3 | 0 | 1 | 0 | 0 | 5 | 
| Total | MOS | 36 | 61 | 59 | 46 | 43 | 6 | 251 | 
|  | CMOS | 38 | 72 | 65 | 58 | 43 | 7 | 283 | 

**Table A1:** Rater pool gender and age distribution

### Inter-rater agreement

In our data, we found a Krippendorff’s alpha for the “Overall” preference question of 0.216, and a Gwet AC2 of 0.211:

| Table A2:  Inter-rater agreement for CMOS metrics (Krippendorff’s alpha, Gwet AC2). In our case (same number of raters per comparison), Krippendorff’s alpha is numerically equivalent to ICC-1 value and hence reported as such.  Low α with high AC2 indicates heavily tie-concentrated distributions. |  |  |  | 
|---|---|---|---|
| CMOS Rubric | α / ICC(1) | Gwet AC2 | n units | 
|---|---|---|---|
| Overall | 0.216 | 0.211 | 5653 | 
| Humanness | 0.234 | 0.495 | 24206 | 
| Conversational register | 0.242 | 0.508 | 5653 | 
| Engagement | 0.207 | 0.459 | 5653 | 
| Pleasantness | 0.234 | 0.501 | 5653 | 
| Naturalness | 0.250 | 0.543 | 5653 | 
| Emotional appropriateness | 0.237 | 0.403 | 1594 | 
| Technical quality | 0.132 | 0.683 | 25175 | 
| Instruction following | 0.225 | 0.538 | 1010 | 
| Perceived audio quality | 0.057 | 0.669 | 5653 | 
| Pronunciation accuracy | 0.054 | 0.880 | 5653 | 
| Voice consistency | 0.047 | 0.711 | 5653 | 
| Response quality | 0.207 | 0.501 | 5653 | 
| Reasoning | 0.184 | 0.577 | 1553 | 

**Table A2:** Inter-rater agreement for CMOS metrics (Krippendorff’s alpha, Gwet AC2). In our case (same number of raters per comparison), Krippendorff’s alpha is numerically equivalent to ICC-1 value and hence reported as such. Low α with high AC2 indicates heavily tie-concentrated distributions.

However, on average each sample pair has only ~2 independent ratings, making traditional inter-rater agreement metrics difficult to interpret. Instead, we prefer split-half Spearman correlation of overall model ranking, which measures how similar the model rankings are when computed from two disjoint random halves of the rater pool. The table below shows the mean correlation of split-half rankings over 20,000 random splits:

| Table A3:  Split-half Spearman correlation for our rankings (CMOS) and standard deviation |  |  |  | 
|---|---|---|---|
| Family | Rubric | ρ | sd | 
|---|---|---|---|
| Overall | Overall preference | 0.964 | 0.013 | 
| Humanness | Emotional appropriateness (gated) | 0.961 | 0.018 | 
|  | Naturalness | 0.955 | 0.023 | 
|  | Engagement | 0.934 | 0.028 | 
|  | Conversational register | 0.923 | 0.026 | 
|  | Pleasantness | 0.917 | 0.029 | 
| Technical Quality | Response quality | 0.938 | 0.028 | 
|  | Perceived audio quality | 0.925 | 0.042 | 
|  | Reasoning (gated) | 0.867 | 0.031 | 
|  | Pronunciation accuracy | 0.813 | 0.073 | 
|  | Instruction following (gated) | 0.802 | 0.069 | 
|  | Voice consistency | 0.674 | 0.122 | 

**Table A3:** Split-half Spearman correlation for our rankings (CMOS) and standard deviation

Together, these metrics tell us that per-rating agreement is highly variable across raters, but in aggregate the overall model ranking produced is very stable.  More discussion of this phenomenon can be found in the appendix section of [our previous blog post](https://research.withdavid.ai/blog/lalm-as-judge-vs-hitl).

### MOS vs. CMOS Agreement

Our published leaderboard relies on CMOS ratings data to drive Elo scores, but we also collected human MOS ratings on the same set of rubrics on a distinct set of input prompts.

| Table A4:  CMOS vs. MOS (Overall) with 95%-CI per model; MOS score is calculated by averaging MOS scores for each sample and averaging all samples. |  |  | 
|---|---|---|
| Model | CMOS Elo | MOS (sample-averaged) | 
|---|---|---|
| GPT-Realtime 2.1 | 1149 [1130, 1168] | 4.13 [4.03, 4.23] | 
| GPT-Realtime 2.0 | 1128 [1109, 1149] | 4.08 [3.98, 4.18] | 
| Gemini 3.1 Flash Live | 1071 [1050, 1091] | 4.26 [4.19, 4.32] | 
| Grok voice-think-fast 2.0 | 993 [977, 1009] | 3.64 [3.53, 3.75] | 
| Nova 2 Sonic | 979 [948, 1014] | 3.59 [3.46, 3.73] | 
| Cascade | 897 [864, 928] | 3.62 [3.47, 3.74] | 
| Qwen3-Omni-30B-A3B-Instruct | 784 [751, 817] | 3.04 [2.91, 3.18] | 

**Table A4:** CMOS vs. MOS (Overall) with 95%-CI per model; MOS score is calculated by averaging MOS scores for each sample and averaging all samples.

While these two metrics largely agree, there are two notable differences:

1. MOS flips the Gemini 3.1 Flash Live and GPT-Realtime-2.0/2.1 rankings at the top.
2. MOS considered the cascade baseline to be competitive with Grok and Nova 2 Sonic, while CMOS considered the cascade baseline to lag far behind.

There are several plausible explanations for why you might see a difference between MOS and CMOS ratings that are worth examining.

#### Case 1: MOS penalizes catastrophic failures more harshly in saturated domains

The acoustic quality of modern TTS systems is high enough to saturate MOS ratings - almost everything is “pretty good” on an objective scale, and model response that’s in the lowest quartile of output quality might still get a 4 on the MOS scale. In contrast, CMOS asks about relative preference, and the extreme values of the scale simply indicate a “strong” subjective preference for one output over another. A model might have a perfectly good output, but a competing model might still be strongly preferred by the rater. This means that a catastrophic failure earns -3 on CMOS, but so does a merely-strongly-dispreferred output, so CMOS caps the penalty of catastrophic samples while MOS doesn't.

The implication is that model A could be preferred over model B in general, but also have more catastrophic failures, which would lead to model A having a higher CMOS rating but a lower MOS rating. In the case of Gemini 3.1 Flash Live and GPT-Realtime-2.1, we see the following distributions:

| Table A5:  Distribution of “overall” MOS ratings for each model |  |  |  |  |  | 
|---|---|---|---|---|---|
|  | 1 | 2 | 3 | 4 | 5 | 
|---|---|---|---|---|---|
| Gemini 3.1 Flash Live | 0.9% | 2.1% | 11.8% | 41% | 44.2% | 
| GPT-Realtime-2.1 | 2.2% | 3.2% | 14.5% | 40.3% | 39.8% | 

**Table A5:** Distribution of “overall” MOS ratings for each model

In our data, it appears that this effect partially explains the MOS / CMOS discrepancy, but does not completely resolve the question. It is true that GPT-Realtime-2.1 MOS shows more catastrophic failures (MOS scores of 1 and 2) than Gemini 3.1 Flash Live, but this is part of an overall left-shift across the whole distribution; there are also more 3’s and fewer 4’s and 5’s.

#### Case 2: Heterogeneity in model performance across the input distribution impacts MOS and CMOS asymmetrically

If one model performs characteristically differently on a subset of input prompts, it could have a similar distributional effect to the catastrophic failure case discussed above; model A might be significantly worse at a subset of the eval, but marginally better at the majority, which would lead to stronger CMOS rankings and weaker MOS rankings. To explore this in our data, we examine whether MOS score difference is specific to the difficulty of the prompt:

| Table A6:  MOS values for GPT-Realtime-2.1 and Gemini 3.1 Flash Live per Difficulty Level (L1 - L5 + degraded category). |  |  |  |  | 
|---|---|---|---|---|
| Difficulty | GPT-Realtime 2.1 | Gemini 3.1 Flash Live | gap | n (GPT-Realtime 2.1 / Gemini) | 
|---|---|---|---|---|
| L1 | 4.16 | 4.30 | +0.13 | 770 / 1140 | 
| L2 | 4.16 | 4.29 | +0.13 | 809 / 1168 | 
| L3 | 4.15 | 4.25 | +0.10 | 844 / 1162 | 
| L4 | 4.07 | 4.21 | +0.14 | 1223 / 1222 | 
| L5 | 4.13 | 4.23 | +0.11 | 1075 / 1076 | 
| Degraded | 4.08 | 4.24 | +0.16 | 349 / 348 | 

**Table A6:** MOS values for GPT-Realtime-2.1 and Gemini 3.1 Flash Live per Difficulty Level (L1 - L5 + degraded category).

As expected, the overall MOS score is marginally higher for easier prompts and lower for harder prompts, but the gap between models remains relatively consistent across difficulty levels.

#### Case 3: Certain dimensions of quality impact overall quality score differently than they impact overall preference

For example, if a rater strongly prefers one speaking voice over another, this could result in a high overall MOS ratings for model A while still showing a strong overall preference for model B on overall CMOS preference. To investigate, we examine the relationship between the MOS / CMOS delta between models across rubric dimensions:

| Table A7:  MOS vs. CMOS per metrics for GPT-Realtime 2.1 - Gemini 3.1 Flash Live. The gated metrics are only applicable to certain prompts. CMOS value of the average head-to-head CMOS score on a scale of (-3..+3) |  |  | 
|---|---|---|
| Rubric | MOS Δ (GPT 2.1 − Gemini) | CMOS Δ (GPT 2.1 vs Gemini) | 
|---|---|---|
| Overall preference | −0.13 | +0.58 | 
| Humanness |  |  | 
| Conversational register | −0.09 | +0.23 | 
| Engagement | −0.26 | +0.18 | 
| Emotional appropriateness (gated) | +0.06 | +0.65 | 
| Naturalness | −0.07 | +0.29 | 
| Pleasantness | −0.00 | +0.35 | 
| Technical quality |  |  | 
| Response quality | +0.06 | +0.60 | 
| Reasoning (gated) | +0.02 | +0.70 | 
| Instruction following (gated) | +0.04 | +0.38 | 
| Perceived audio quality | −0.19 | −0.15 | 
| Pronunciation accuracy | −0.11 | −0.04 | 
| Voice consistency | −0.17 | +0.04 | 

**Table A7:** MOS vs. CMOS per metrics for GPT-Realtime 2.1 - Gemini 3.1 Flash Live. The gated metrics are only applicable to certain prompts. CMOS value of the average head-to-head CMOS score on a scale of (-3..+3)

Here again, we find the overall MOS ratings favoring Gemini 3.1 Flash Live across most dimensions, while the CMOS ratings favor GPT-Realtime-2.1 across most dimensions. Interestingly, we do see that perceived audio quality favors Gemini for both MOS and CMOS. One interpretation of this might be that in isolation raters tended to penalize GPT-Realtime for audio quality issues, but this factor did not strongly influence their overall preference across model outputs.

Our conclusion is that MOS and CMOS are simply measuring different things; it is entirely plausible that there are some model quality shortcomings that are relatively more obvious in a standalone pointwise setting, and others that are relatively more obvious in a comparative setting.

### Rubric clustering and correlation

Each rubric category bundles several sub-metrics which are themselves positively correlated across prompts within a cluster. Humanness is highly coherent: its five delivery/affect sub-questions (naturalness, pleasantness, engagement, emotional appropriateness, conversational register) move together at a mean pairwise correlation of 0.68, with naturalness and pleasantness the tightest pair (0.82).

Technical Quality is more loosely coupled, with a mean correlation of 0.40. One explanation for this is that it spans two distinct sub-groups: the lexical content items — response quality, instruction following and reasoning — track each other closely at 0.82, while the acoustic content items — perceived audio quality, pronunciation accuracy and voice consistency — cohere more modestly at 0.47. These two sub-groups of Technical Quality correlate only weakly with each other (0.28).

| Table A8:  Metrics correlation within a category group and across category groups |  |  | 
|---|---|---|
| Group / cluster | Rubrics | Mean pairwise ρ | 
|---|---|---|
| Humanness | naturalness, pleasantness, engagement, emotional appropriateness, conversational register | 0.68 | 
| Technical Quality | response quality, instruction following, reasoning, perceived audio quality, pronunciation accuracy, voice consistency | 0.40 | 
| Technical Quality - lexical | response quality, instruction following, reasoning | 0.82 | 
| Technical Quality - audio | perceived audio quality, pronunciation accuracy, voice consistency | 0.47 | 
| lexical × audio sub-group | cross-group | 0.28 | 

**Table A8:** Metrics correlation within a category group and across category groups

Humanness can be interpreted as a single tight group, while Technical Quality is a broader axis holding a strong lexical content block and a looser audio content block.

These two sub-groups are exactly the groupings that emerge for a k=3 clustering of our data, so it is worth saying why we do not separate them out: we do not find this third cluster to reproduce reliably across rater-prompt bootstrapping (see methodology section). We hypothesize that with more evaluation data and more models being evaluated we would see this third grouping emerge reliably, but our data as yet does not justify treating these as separate clusters, so for the time being we continue using the k=2 grouping of Humanness and Technical Quality that robustly reproduces.

### Insights from free text analysis

Our collection protocol includes the ability for raters to leave free text comments explaining their ratings, rationale, or any issues they want to flag that are not captured in existing rubrics. We perform clustering analysis around themes from free-text rater comments and display the results in the table below. Theme mention rate is calculated as percentage of the model’s rater comments from the MOS ratings, where (as discussed in [previous post](https://research.withdavid.ai/blog/lalm-as-judge-vs-hitl)) we see more robust and informative free text comments. Note that this analysis shows the relative differences between the way raters tend to characterize the responses given by each model, rather than the absolute number of comments that addressed the particular theme. Our hope is that these qualitative insights are informative to model trainers in identifying areas of strength and weakness of their model, and that they may influence future model training priorities. 

| Table A9:  Theme mention rate by model, as a percentage of each model’s MOS rater comments. Shading compares across a row, not down a column. Themes overlap, so rates do not sum to 100. |  |  |  |  |  |  |  | 
|---|---|---|---|---|---|---|---|
| Theme | GPT-Realtime 2.1 | GPT-Realtime 2.0 | Gemini 3.1 Flash | Grok voice-think-fast | Nova 2 Sonic | Cascade | Qwen3-Omni | 
|---|---|---|---|---|---|---|---|
| Natural / human | 17.0 | 15.4 | 17.0 | 10.5 | 9.2 | 10.7 | 8.5 | 
| Robotic / monotone | 5.4 | 7.7 | 4.4 | 19.5 | 16.9 | 23.6 | 15.2 | 
| Warm / pleasant | 9.1 | 8.0 | 8.5 | 6.0 | 6.1 | 4.4 | 4.1 | 
| Helpful / informative | 33.9 | 33.7 | 34.5 | 30.4 | 29.2 | 37.2 | 21.3 | 
| Too long / verbose | 2.8 | 2.7 | 0.6 | 2.1 | 14.3 | 0.8 | 0.5 | 
| Pacing (fast / slow) | 5.5 | 5.3 | 2.8 | 3.6 | 5.5 | 8.2 | 14.6 | 
| Audio artifact | 11.6 | 10.8 | 5.8 | 13.3 | 24.1 | 8.0 | 15.2 | 
| Accent / mispronunciation | 4.9 | 5.4 | 3.1 | 4.4 | 2.9 | 5.3 | 10.3 | 
| Emotional / empathetic | 6.3 | 5.2 | 5.8 | 3.2 | 2.1 | 2.7 | 2.5 | 
| Breathy / breathing | 1.9 | 1.8 | 5.2 | 0.7 | 0.9 | 0.1 | 1.6 | 

**Table A9:** Theme mention rate by model, as a percentage of each model’s MOS rater comments. Shading compares across a row, not down a column. Themes overlap, so rates do not sum to 100.

### Rubric-Level MOS & CMOS Rankings

Below is the full list of CMOS Elo ranking and MOS scores per model per metric with 95% CI; the MOS score is calculated using sample-averaging. CMOS Elo rankings follow the same BT fitting approach discussed in the methodology section, simply using the specific rubric preference question rather than the overall preference question.

| Table A10:  CMOS Elo and sample-averaged MOS per model for every rubric, with 95% CIs in brackets. |  |  | 
|---|---|---|
| Family | CMOS Elo | MOS 1–5 (sample-averaged) | 
|---|---|---|
| Overall preference |  |  | 
| GPT-Realtime 2.1 | 1149 [1130, 1168] | 4.13 [4.03, 4.23] | 
| GPT-Realtime 2.0 | 1128 [1109, 1149] | 4.08 [3.98, 4.18] | 
| Gemini 3.1 Flash Live | 1071 [1050, 1091] | 4.26 [4.19, 4.32] | 
| Grok voice-think-fast 2.0 | 993 [977, 1009] | 3.64 [3.53, 3.75] | 
| Nova 2 Sonic | 979 [948, 1014] | 3.59 [3.46, 3.73] | 
| Cascade | 897 [864, 928] | 3.62 [3.47, 3.74] | 
| Qwen3-Omni-30B-A3B-Instruct | 784 [751, 817] | 3.04 [2.91, 3.18] | 
| Humanness |  |  | 
| Conversational register |  |  | 
| GPT-Realtime 2.1 | 1160 [1140, 1179] | 4.24 [4.16, 4.33] | 
| GPT-Realtime 2.0 | 1131 [1109, 1152] | 4.17 [4.07, 4.26] | 
| Gemini 3.1 Flash Live | 1122 [1100, 1143] | 4.33 [4.25, 4.42] | 
| Grok voice-think-fast 2.0 | 981 [961, 1002] | 3.71 [3.58, 3.83] | 
| Nova 2 Sonic | 921 [888, 956] | 3.59 [3.44, 3.73] | 
| Cascade | 844 [808, 875] | 3.44 [3.26, 3.61] | 
| Qwen3-Omni-30B-A3B-Instruct | 841 [798, 880] | 3.14 [2.95, 3.32] | 
| Engagement |  |  | 
| GPT-Realtime 2.1 | 1138 [1118, 1155] | 4.00 [3.91, 4.10] | 
| Gemini 3.1 Flash Live | 1105 [1080, 1133] | 4.25 [4.18, 4.33] | 
| GPT-Realtime 2.0 | 1102 [1080, 1120] | 3.90 [3.78, 4.01] | 
| Nova 2 Sonic | 988 [957, 1029] | 3.46 [3.31, 3.60] | 
| Grok voice-think-fast 2.0 | 971 [952, 990] | 3.54 [3.42, 3.65] | 
| Cascade | 883 [853, 912] | 3.41 [3.25, 3.57] | 
| Qwen3-Omni-30B-A3B-Instruct | 813 [782, 847] | 3.12 [2.96, 3.26] | 
| Emotional appropriateness |  |  | 
| GPT-Realtime 2.1 | 1179 [1157, 1208] | 4.43 [4.35, 4.52] | 
| GPT-Realtime 2.0 | 1138 [1104, 1173] | 4.36 [4.24, 4.47] | 
| Gemini 3.1 Flash Live | 1073 [1044, 1105] | 4.37 [4.28, 4.46] | 
| Grok voice-think-fast 2.0 | 977 [950, 1008] | 3.77 [3.64, 3.91] | 
| Nova 2 Sonic | 964 [921, 1003] | 3.98 [3.81, 4.16] | 
| Cascade | 866 [811, 916] | 3.68 [3.48, 3.86] | 
| Qwen3-Omni-30B-A3B-Instruct | 802 [747, 847] | 3.14 [2.94, 3.35] | 
| Naturalness |  |  | 
| GPT-Realtime 2.1 | 1176 [1160, 1196] | 4.14 [4.06, 4.24] | 
| GPT-Realtime 2.0 | 1149 [1128, 1170] | 4.10 [4.01, 4.20] | 
| Gemini 3.1 Flash Live | 1127 [1106, 1147] | 4.22 [4.13, 4.30] | 
| Grok voice-think-fast 2.0 | 977 [954, 998] | 3.52 [3.41, 3.65] | 
| Nova 2 Sonic | 960 [935, 988] | 3.60 [3.47, 3.73] | 
| Cascade | 810 [771, 844] | 3.11 [2.90, 3.30] | 
| Qwen3-Omni-30B-A3B-Instruct | 800 [759, 840] | 2.78 [2.62, 2.93] | 
| Pleasantness |  |  | 
| GPT-Realtime 2.1 | 1164 [1146, 1183] | 4.18 [4.09, 4.27] | 
| GPT-Realtime 2.0 | 1134 [1115, 1155] | 4.14 [4.04, 4.24] | 
| Gemini 3.1 Flash Live | 1107 [1085, 1132] | 4.18 [4.10, 4.26] | 
| Grok voice-think-fast 2.0 | 987 [966, 1005] | 3.57 [3.45, 3.69] | 
| Nova 2 Sonic | 978 [950, 1003] | 3.73 [3.61, 3.86] | 
| Cascade | 847 [809, 885] | 3.39 [3.23, 3.56] | 
| Qwen3-Omni-30B-A3B-Instruct | 783 [748, 821] | 2.94 [2.79, 3.08] | 
| Technical quality |  |  | 
| Response quality |  |  | 
| GPT-Realtime 2.1 | 1113 [1099, 1128] | 4.39 [4.33, 4.46] | 
| GPT-Realtime 2.0 | 1105 [1087, 1125] | 4.41 [4.33, 4.48] | 
| Nova 2 Sonic | 1050 [1020, 1080] | 4.14 [4.04, 4.24] | 
| Gemini 3.1 Flash Live | 1009 [991, 1027] | 4.33 [4.26, 4.40] | 
| Grok voice-think-fast 2.0 | 1006 [990, 1021] | 4.19 [4.12, 4.27] | 
| Cascade | 947 [918, 974] | 4.16 [4.06, 4.26] | 
| Qwen3-Omni-30B-A3B-Instruct | 770 [742, 798] | 3.71 [3.60, 3.82] | 
| Instruction following |  |  | 
| GPT-Realtime 2.1 | 1075 [1038, 1106] | 4.52 [4.43, 4.61] | 
| GPT-Realtime 2.0 | 1064 [1023, 1103] | 4.53 [4.41, 4.64] | 
| Gemini 3.1 Flash Live | 1015 [982, 1049] | 4.49 [4.39, 4.57] | 
| Grok voice-think-fast 2.0 | 1015 [979, 1048] | 4.41 [4.30, 4.51] | 
| Cascade | 1003 [940, 1083] | 4.39 [4.22, 4.55] | 
| Nova 2 Sonic | 966 [902, 1026] | 4.33 [4.16, 4.48] | 
| Qwen3-Omni-30B-A3B-Instruct | 861 [815, 914] | 4.00 [3.79, 4.19] | 
| Reasoning |  |  | 
| GPT-Realtime 2.1 | 1118 [1095, 1145] | 4.70 [4.63, 4.76] | 
| Nova 2 Sonic | 1109 [1055, 1162] | 4.47 [4.35, 4.59] | 
| GPT-Realtime 2.0 | 1103 [1073, 1137] | 4.70 [4.62, 4.78] | 
| Grok voice-think-fast 2.0 | 1007 [978, 1037] | 4.55 [4.47, 4.64] | 
| Gemini 3.1 Flash Live | 979 [949, 1008] | 4.68 [4.62, 4.74] | 
| Cascade | 927 [870, 981] | 4.45 [4.33, 4.57] | 
| Qwen3-Omni-30B-A3B-Instruct | 757 [700, 809] | 4.16 [4.00, 4.30] | 
| Perceived audio quality |  |  | 
| Gemini 3.1 Flash Live | 1103 [1082, 1126] | 4.51 [4.45, 4.58] | 
| GPT-Realtime 2.1 | 1071 [1048, 1093] | 4.32 [4.23, 4.42] | 
| GPT-Realtime 2.0 | 1068 [1047, 1090] | 4.31 [4.22, 4.39] | 
| Grok voice-think-fast 2.0 | 1015 [993, 1039] | 4.16 [4.04, 4.27] | 
| Cascade | 990 [947, 1033] | 4.17 [4.04, 4.29] | 
| Nova 2 Sonic | 915 [883, 949] | 3.91 [3.77, 4.05] | 
| Qwen3-Omni-30B-A3B-Instruct | 837 [796, 873] | 3.68 [3.48, 3.86] | 
| Pronunciation accuracy |  |  | 
| Gemini 3.1 Flash Live | 1073 [1053, 1094] | 4.76 [4.71, 4.81] | 
| GPT-Realtime 2.0 | 1064 [1038, 1087] | 4.63 [4.57, 4.70] | 
| GPT-Realtime 2.1 | 1056 [1035, 1079] | 4.65 [4.58, 4.71] | 
| Nova 2 Sonic | 1026 [1000, 1048] | 4.64 [4.56, 4.72] | 
| Grok voice-think-fast 2.0 | 1018 [996, 1040] | 4.64 [4.58, 4.70] | 
| Cascade | 1002 [964, 1039] | 4.55 [4.45, 4.63] | 
| Qwen3-Omni-30B-A3B-Instruct | 762 [716, 802] | 4.01 [3.87, 4.14] | 
| Voice consistency |  |  | 
| GPT-Realtime 2.0 | 1066 [1041, 1087] | 4.53 [4.45, 4.62] | 
| GPT-Realtime 2.1 | 1040 [1015, 1062] | 4.51 [4.42, 4.60] | 
| Gemini 3.1 Flash Live | 1039 [1013, 1068] | 4.68 [4.62, 4.73] | 
| Grok voice-think-fast 2.0 | 1023 [999, 1044] | 4.40 [4.32, 4.49] | 
| Nova 2 Sonic | 1020 [992, 1048] | 4.55 [4.47, 4.63] | 
| Cascade | 985 [949, 1022] | 4.35 [4.23, 4.47] | 
| Qwen3-Omni-30B-A3B-Instruct | 827 [781, 875] | 4.04 [3.90, 4.16] | 

**Table A10:** CMOS Elo and sample-averaged MOS per model for every rubric, with 95% CIs in brackets.

## **Model Configuration Details**

## OpenAIGPT-Realtime-2.0-xhigh (Marin)

- Config (simplified)
- Model: GPT-Realtime 2.0 Voice: Marin Reasoning: xhigh Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "gpt-realtime-2" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "marin" } } reasoning: { effort: "xhigh" } instructions: "You are a helpful voice assistant."

## OpenAIGPT-Realtime-2.0-xhigh (Cedar)

- Config (simplified)
- Model: GPT-Realtime 2.0 Voice: Cedar Reasoning: xhigh Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "gpt-realtime-2" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "cedar" } } reasoning: { effort: "xhigh" } instructions: "You are a helpful voice assistant."

## OpenAIGPT-Realtime-2.1-xhigh (Marin)

- Config (simplified)
- Model: GPT-Realtime 2.1 Voice: Marin Reasoning: xhigh Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "gpt-realtime-2.1" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "marin" } } reasoning: { effort: "xhigh" } instructions: "You are a helpful voice assistant."

## OpenAIGPT-Realtime-2.1-xhigh (Cedar)

- Config (simplified)
- Model: GPT-Realtime 2.1 Voice: Cedar Reasoning: xhigh Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "gpt-realtime-2.1" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "cedar" } } reasoning: { effort: "xhigh" } instructions: "You are a helpful voice assistant."

## OpenAIGPT-Realtime-2.1-minimal (Marin)

- Config (simplified)
- Model: GPT-Realtime 2.1 Voice: Marin Reasoning: minimal Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "gpt-realtime-2.1" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "marin" } } reasoning: { effort: "minimal" } instructions: "You are a helpful voice assistant."

## OpenAIGPT-Realtime-2.1-minimal (Cedar)

- Config (simplified)
- Model: GPT-Realtime 2.1 Voice: Cedar Reasoning: minimal Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "gpt-realtime-2.1" type: "realtime" output_modalities: ["audio"] audio: { input: { format: { type: "audio/pcm", rate: 24000 }, turn_detection: null }, output: { format: { type: "audio/pcm", rate: 24000 }, voice: "cedar" } } reasoning: { effort: "minimal" } instructions: "You are a helpful voice assistant."

## xAIGrok Voice Think Fast 2.0 - High (Ara)

- Config (simplified)
- Model: Grok Voice Think Fast 2.0 Voice: Ara Reasoning: high Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "grok-voice-think-fast-2.0" session: { instructions: "You are a helpful voice assistant.", voice: "ara", reasoning: { effort: "high" }, turn_detection: { type: null }, audio: { input: { format: { type: "audio/pcm", rate: 24000 } }, output: { format: { type: "audio/pcm", rate: 24000 } } } } Supported voices: ara, eve, leo, rex, sal

## xAIGrok Voice Think Fast 2.0 - High (Rex)

- Config (simplified)
- Model: Grok Voice Think Fast 2.0 Voice: Rex Reasoning: high Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "grok-voice-think-fast-2.0" session: { instructions: "You are a helpful voice assistant.", voice: "rex", reasoning: { effort: "high" }, turn_detection: { type: null }, audio: { input: { format: { type: "audio/pcm", rate: 24000 } }, output: { format: { type: "audio/pcm", rate: 24000 } } } } Supported voices: ara, eve, leo, rex, sal

## xAIGrok Voice Think Fast 2.0 - None (Ara)

- Config (simplified)
- Model: Grok Voice Think Fast 2.0 Voice: Ara Reasoning: none (xAI's own default is high, so this is sent explicitly) Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "grok-voice-think-fast-2.0" session: { instructions: "You are a helpful voice assistant.", voice: "ara", reasoning: { effort: "none" }, turn_detection: { type: null }, audio: { input: { format: { type: "audio/pcm", rate: 24000 } }, output: { format: { type: "audio/pcm", rate: 24000 } } } } Supported voices: ara, eve, leo, rex, sal

## xAIGrok Voice Think Fast 2.0 - None (Rex)

- Config (simplified)
- Model: Grok Voice Think Fast 2.0 Voice: Rex Reasoning: none (xAI's own default is high, so this is sent explicitly) Audio: PCM16, 24 kHz in / 24 kHz out Turn detection: off — whole clip sent, then one reply Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "grok-voice-think-fast-2.0" session: { instructions: "You are a helpful voice assistant.", voice: "rex", reasoning: { effort: "none" }, turn_detection: { type: null }, audio: { input: { format: { type: "audio/pcm", rate: 24000 } }, output: { format: { type: "audio/pcm", rate: 24000 } } } } Supported voices: ara, eve, leo, rex, sal

## GoogleGemini 3.1 Flash Live Preview - thinking-high (Zephyr)

- Config (simplified)
- Model: Gemini 3.1 Flash Live Preview Voice: Zephyr Thinking: high — thoughts not returned Audio: PCM16, 16 kHz in / 24 kHz out Turn detection: off — automatic activity detection disabled Extras: output audio transcription on Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "gemini-3.1-flash-live-preview" responseModalities: ["AUDIO"] systemInstruction: "You are a helpful voice assistant." speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: "Zephyr" } } } realtimeInputConfig: { automaticActivityDetection: { disabled: true } } thinkingConfig: { thinkingLevel: "HIGH", includeThoughts: false } outputAudioTranscription: {} realtimeInput: { audio: { data: <b64>, mimeType: "audio/pcm;rate=16000" } } Output audio: 24000 Hz PCM16

## GoogleGemini 3.1 Flash Live Preview - thinking-high (Iapetus)

- Config (simplified)
- Model: Gemini 3.1 Flash Live Preview Voice: Iapetus Thinking: high — thoughts not returned Audio: PCM16, 16 kHz in / 24 kHz out Turn detection: off — automatic activity detection disabled Extras: output audio transcription on Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "gemini-3.1-flash-live-preview" responseModalities: ["AUDIO"] systemInstruction: "You are a helpful voice assistant." speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: "Iapetus" } } } realtimeInputConfig: { automaticActivityDetection: { disabled: true } } thinkingConfig: { thinkingLevel: "HIGH", includeThoughts: false } outputAudioTranscription: {} realtimeInput: { audio: { data: <b64>, mimeType: "audio/pcm;rate=16000" } } Output audio: 24000 Hz PCM16

## GoogleGemini 3.1 Flash Live Preview - thinking-minimal (Zephyr)

- Config (simplified)
- Model: Gemini 3.1 Flash Live Preview Voice: Zephyr Thinking: minimal — thoughts not returned Audio: PCM16, 16 kHz in / 24 kHz out Turn detection: off — automatic activity detection disabled Extras: output audio transcription on Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "gemini-3.1-flash-live-preview" responseModalities: ["AUDIO"] systemInstruction: "You are a helpful voice assistant." speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: "Zephyr" } } } realtimeInputConfig: { automaticActivityDetection: { disabled: true } } thinkingConfig: { thinkingLevel: "MINIMAL", includeThoughts: false } outputAudioTranscription: {} realtimeInput: { audio: { data: <b64>, mimeType: "audio/pcm;rate=16000" } } Output audio: 24000 Hz PCM16

## GoogleGemini 3.1 Flash Live Preview - thinking-minimal (Iapetus)

- Config (simplified)
- Model: Gemini 3.1 Flash Live Preview Voice: Iapetus Thinking: minimal — thoughts not returned Audio: PCM16, 16 kHz in / 24 kHz out Turn detection: off — automatic activity detection disabled Extras: output audio transcription on Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "gemini-3.1-flash-live-preview" responseModalities: ["AUDIO"] systemInstruction: "You are a helpful voice assistant." speechConfig: { voiceConfig: { prebuiltVoiceConfig: { voiceName: "Iapetus" } } } realtimeInputConfig: { automaticActivityDetection: { disabled: true } } thinkingConfig: { thinkingLevel: "MINIMAL", includeThoughts: false } outputAudioTranscription: {} realtimeInput: { audio: { data: <b64>, mimeType: "audio/pcm;rate=16000" } } Output audio: 24000 Hz PCM16

## AmazonNova 2 Sonic (Tiffany)

- Config (simplified)
- Model: Nova 2 Sonic (Bedrock, us-east-1) Voice: Tiffany Reasoning: n/a — no effort control Audio: LPCM16 mono, 16 kHz in / 24 kHz out Turn detection: cannot be disabled — endpointing LOW (~2.0 s pause), the most patient setting Sampling: model defaults (no max tokens, top-p or temperature set) Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "amazon.nova-2-sonic-v1:0" (bedrock, us-east-1) sessionStart.inferenceConfiguration: {} # maxTokens / topP / temperature all unset sessionStart.turnDetectionConfiguration: { endpointingSensitivity: "LOW" } promptStart.textOutputConfiguration: { mediaType: "text/plain" } promptStart.audioOutputConfiguration: { mediaType: "audio/lpcm", sampleRateHertz: 24000, sampleSizeBits: 16, channelCount: 1, audioType: "SPEECH", encoding: "base64", voiceId: "tiffany" } contentStart(SYSTEM,TEXT) + textInput.content: "You are a helpful voice assistant." contentStart(USER,AUDIO).audioInputConfiguration: { mediaType: "audio/lpcm", sampleRateHertz: 16000, sampleSizeBits: 16, channelCount: 1, audioType: "SPEECH", encoding: "base64" }

## AmazonNova 2 Sonic (Matthew)

- Config (simplified)
- Model: Nova 2 Sonic (Bedrock, us-east-1) Voice: Matthew Reasoning: n/a — no effort control Audio: LPCM16 mono, 16 kHz in / 24 kHz out Turn detection: cannot be disabled — endpointing LOW (~2.0 s pause), the most patient setting Sampling: model defaults (no max tokens, top-p or temperature set) Prompt: default ("You are a helpful voice assistant.")
- System prompt
- You are a helpful voice assistant.
- Full parameters
- model: "amazon.nova-2-sonic-v1:0" (bedrock, us-east-1) sessionStart.inferenceConfiguration: {} # maxTokens / topP / temperature all unset sessionStart.turnDetectionConfiguration: { endpointingSensitivity: "LOW" } promptStart.textOutputConfiguration: { mediaType: "text/plain" } promptStart.audioOutputConfiguration: { mediaType: "audio/lpcm", sampleRateHertz: 24000, sampleSizeBits: 16, channelCount: 1, audioType: "SPEECH", encoding: "base64", voiceId: "matthew" } contentStart(SYSTEM,TEXT) + textInput.content: "You are a helpful voice assistant." contentStart(USER,AUDIO).audioInputConfiguration: { mediaType: "audio/lpcm", sampleRateHertz: 16000, sampleSizeBits: 16, channelCount: 1, audioType: "SPEECH", encoding: "base64" }

## Alibaba (self-hosted, Modal)Qwen3-Omni-30B-A3B-Instruct (Chelsie)

- Config (simplified)
- Model: Qwen3-Omni 30B-A3B Instruct — self-hosted on Modal (H200) Voice: Chelsie Reasoning: n/a Audio: WAV PCM16, 16 kHz in / 24 kHz out Turn detection: n/a — one clip per call, batched Prompt: Qwen's own official assistant prompt. Caps replies at 50 words, forbids formatting.
- System prompt
- You are a virtual voice assistant with no gender or age. You are communicating with the user. In user messages, "I/me/my/we/our" refer to the user and "you/your" refer to the assistant. In your replies, address the user as "you/your" and yourself as "I/me/my"; never mirror the user's pronouns—always shift perspective. Keep original pronouns only in direct quotes; if a reference is unclear, ask a brief clarifying question. Interact with users using short (no more than 50 words), brief, straightforward language, maintaining a natural tone. Never use formal phrasing, mechanical expressions, bullet points, overly structured language. Your output must consist only of the spoken content you want the user to hear. Do not include any descriptions of actions, emotions, sounds, or voice changes. Do not use asterisks, brackets, parentheses, or any other symbols to indicate tone or actions. You must answer users' audio or text questions, do not directly describe the video content. You should communicate in the same language strictly as the user unless they request otherwise. When you are uncertain (e.g., you can't see/hear clearly, don't understand, or the user makes a comment rather than asking a question), use appropriate questions to guide the user to continue the conversation. Keep replies concise and conversational, as if talking face-to-face
- Full parameters
- model: "Qwen/Qwen3-Omni-30B-A3B-Instruct" (self-hosted) endpoint: modal://main/qwen3-omni-instruct/Qwen3OmniS2SBatched.generate task: "s2s" speaker: "chelsie" messages: [ { role: "system", content: [{ type: "text", text: <official prompt> }] }, { role: "user", content: [{ type: "audio", audio_b64: <WAV b64> }] } ] Input audio: WAV PCM16 16000 Hz Output audio: 24000 Hz PCM16

## Alibaba (self-hosted, Modal)Qwen3-Omni-30B-A3B-Instruct (Ethan)

- Config (simplified)
- Model: Qwen3-Omni 30B-A3B Instruct — self-hosted on Modal (H200) Voice: Ethan Reasoning: n/a Audio: WAV PCM16, 16 kHz in / 24 kHz out Turn detection: n/a — one clip per call, batched Prompt: Qwen's own official assistant prompt. Caps replies at 50 words, forbids formatting.
- System prompt
- You are a virtual voice assistant with no gender or age. You are communicating with the user. In user messages, "I/me/my/we/our" refer to the user and "you/your" refer to the assistant. In your replies, address the user as "you/your" and yourself as "I/me/my"; never mirror the user's pronouns—always shift perspective. Keep original pronouns only in direct quotes; if a reference is unclear, ask a brief clarifying question. Interact with users using short (no more than 50 words), brief, straightforward language, maintaining a natural tone. Never use formal phrasing, mechanical expressions, bullet points, overly structured language. Your output must consist only of the spoken content you want the user to hear. Do not include any descriptions of actions, emotions, sounds, or voice changes. Do not use asterisks, brackets, parentheses, or any other symbols to indicate tone or actions. You must answer users' audio or text questions, do not directly describe the video content. You should communicate in the same language strictly as the user unless they request otherwise. When you are uncertain (e.g., you can't see/hear clearly, don't understand, or the user makes a comment rather than asking a question), use appropriate questions to guide the user to continue the conversation. Keep replies concise and conversational, as if talking face-to-face
- Full parameters
- model: "Qwen/Qwen3-Omni-30B-A3B-Instruct" (self-hosted) endpoint: modal://main/qwen3-omni-instruct/Qwen3OmniS2SBatched.generate task: "s2s" speaker: "ethan" messages: [ { role: "system", content: [{ type: "text", text: <official prompt> }] }, { role: "user", content: [{ type: "audio", audio_b64: <WAV b64> }] } ] Input audio: WAV PCM16 16000 Hz Output audio: 24000 Hz PCM16

## ElevenLabs, OpenAI, ElevenLabsCascade

- Config (simplified)
- Pipeline: ElevenLabs Scribe v2 (ASR) -> OpenAI GPT-5.6-terra (LLM) -> ElevenLabs Flash v2.5 (TTS) Voice: Rachel (ElevenLabs legacy default, retires 2026-12-31) Reasoning: low — on the LLM leg Audio: PCM, 16 kHz in / 24 kHz out Turn detection: n/a — one clip per call Language: auto (not pinned) Prompt: pinned by the deployed app, cannot be overridden. Caps replies at 100 words, forbids markdown.
- System prompt
- You are a helpful assistant handling a voice chat with a user. # Important Voice Considerations 1. Respond naturally and conversationally as you would in a real conversation 2. Try to be helpful and always follow the instructions below. 3. Keep the response short and concise. Do not exceed 100 words. You are writing a final script for text-to-speech (TTS). Your response will be synthesized directly into speech. Follow the duration instruction as strictly as possible. Output only the final spoken text, with natural punctuation. Do not output markdown, bullets, JSON, XML tags, stage directions, or extra commentary. Do not mention these instructions.
- Full parameters
- endpoint: modal://main/s2s-cascaded/SingleTurnCascade.run_turn asr_model: "scribe_v2" (ElevenLabs Scribe v2) llm_model: "gpt-5.6-terra" (OpenAI) tts_provider: "elevenlabs" tts_model: "eleven_flash_v2_5" voice: "21m00Tcm4TlvDq8ikWAM" (Rachel; legacy default, retires 2026-12-31) language: null reasoning_effort: "low" messages: [{ role: "user", content: [{ type: "audio", audio_b64: <WAV b64>, filename: "prompt.wav" }] }] Input audio: 16000 Hz Output audio: 24000 Hz PCM No system-prompt input: the deployed app pins its own and 400s on a system turn.

## **References**

Bradley, R. A., & Terry, M. E. (1952). Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3/4), 324–345.

Chiang, W.-L., et al. (2024). Chatbot Arena: an open platform for evaluating LLMs by human preference.

Elo, A. E. (1978). The Rating of Chessplayers, Past and Present. Arco.

Eyben, F., Scherer, K. R., Schuller, B. W., et al. (2016). The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing. IEEE Transactions on Affective Computing, 7(2), 190–202.

Krippendorff, K. (2004). Content Analysis: An Introduction to Its Methodology (2nd ed.). Sage.

Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29–48.

McCullagh, P. (1980). Regression models for ordinal data. JRSS: Series B, 42(2), 109–142. (cumulative-link / ordinal model)

Massey, K. (1997). Statistical models applied to the rating of sports teams. Bluefield College.

Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428.
