{"slug": "asymmetric-classifier-free-guidance-for-target-speaker-asr", "title": "Asymmetric Classifier-Free Guidance for Target-Speaker ASR", "summary": "Researchers introduced asymmetric classifier-free guidance (CFG) for target-speaker automatic speech recognition using Whisper, reporting relative word error rate reductions of up to 21.8% over the condition-only baseline and 5.6% over standard conditional decoding of the same CFG-trained model under domain shifts. The method uses a speaker-conditioned branch to predict the target transcript and a speaker-unconditioned branch to predict serialized multi-speaker transcripts, with a global guidance scale selected on target-domain development data and a lightweight encoder-based predictor adjusting it per utterance while the recognition model stays fixed. Oracle analysis indicates substantially larger WER reductions are possible through utterance-level scale selection.", "body_md": "arXiv:2609.30476v1 Announce Type: cross \nAbstract: Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized multi-speaker transcripts. CFG adjusts the contribution of speaker conditioning during decoding through a single guidance scale. We select a global guidance scale on target-domain development data and train a lightweight encoder-based predictor to adjust it for each utterance, keeping the recognition model fixed. Under domain shifts, our full system achieves relative word error rate (WER) reductions of up to 21.8% over the condition-only baseline, and 5.6% over standard conditional decoding of the same CFG-trained model. Oracle analysis shows that substantially larger WER reductions are possible through utterance-level scale selection and identifies how beneficial adjustments vary with domain shifts.", "url": "https://wpnews.pro/news/asymmetric-classifier-free-guidance-for-target-speaker-asr", "canonical_source": "https://www.machinebrief.com/news/asymmetric-classifier-free-guidance-for-target-speaker-asr-mmrg", "published_at": "2026-09-28 04:00:00+00:00", "updated_at": "2026-09-28 05:48:09.108553+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "ai-research"], "entities": ["Whisper", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/asymmetric-classifier-free-guidance-for-target-speaker-asr", "markdown": "https://wpnews.pro/news/asymmetric-classifier-free-guidance-for-target-speaker-asr.md", "text": "https://wpnews.pro/news/asymmetric-classifier-free-guidance-for-target-speaker-asr.txt", "jsonld": "https://wpnews.pro/news/asymmetric-classifier-free-guidance-for-target-speaker-asr.jsonld"}}