{"slug": "does-a-prosody-trained-representation-help-beyond-trainable-fusion-a-parameter", "title": "Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT", "summary": "A parameter-matched study using a frozen HuBERT backbone found that a 64-dimensional prosody representation trained to predict log F0, voicing, Delta log F0, log energy, and spectral tilt produced no measurable incremental word error rate (WER) benefit over a trainable fusion control with zero auxiliary input. Across Buckeye, Switchboard, and AMI IHM, the Null fusion condition reduced WER by 0.71-1.45 points over the frozen-backbone Baseline, while the Learned condition differed from Null by only +0.07, -0.09, and +0.00 points, with no significant differences. Removing or mismatching the representation at inference increased Learned WER, indicating the model depends on the representation despite the absence of measurable incremental gains.", "body_md": "arXiv:2609.36754v1 Announce Type: cross \nAbstract: Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBERT backbone and a 64-dimensional representation trained to predict log F0, voicing, Delta log F0, log energy, and spectral tilt. We compare a frozen-backbone recognizer (Baseline), trainable fusion with zero auxiliary input (Null), and the same fusion supplied with the learned representation (Learned). Across Buckeye, Switchboard, and AMI IHM, Null reduces WER by 0.71-1.45 points over Baseline, whereas Learned differs from Null by +0.07, -0.09, and +0.00 points, with no significant differences. However, removing or mismatching the representation at inference increases Learned WER. Thus, Learned depends on the representation yet shows no measurable incremental WER benefit over the parameter-matched control.", "url": "https://wpnews.pro/news/does-a-prosody-trained-representation-help-beyond-trainable-fusion-a-parameter", "canonical_source": "https://www.machinebrief.com/news/does-a-prosody-trained-representation-help-beyond-trainable-j6k5", "published_at": "2026-09-30 04:00:00+00:00", "updated_at": "2026-09-30 05:47:17.481319+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "ai-research"], "entities": ["HuBERT", "Buckeye", "Switchboard", "AMI IHM", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/does-a-prosody-trained-representation-help-beyond-trainable-fusion-a-parameter", "markdown": "https://wpnews.pro/news/does-a-prosody-trained-representation-help-beyond-trainable-fusion-a-parameter.md", "text": "https://wpnews.pro/news/does-a-prosody-trained-representation-help-beyond-trainable-fusion-a-parameter.txt", "jsonld": "https://wpnews.pro/news/does-a-prosody-trained-representation-help-beyond-trainable-fusion-a-parameter.jsonld"}}