arXiv:2609.30476v1 Announce Type: cross Abstract: Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized multi-speaker transcripts. CFG adjusts the contribution of speaker conditioning during decoding through a single guidance scale. We select a global guidance scale on target-domain development data and train a lightweight encoder-based predictor to adjust it for each utterance, keeping the recognition model fixed. Under domain shifts, our full system achieves relative word error rate (WER) reductions of up to 21.8% over the condition-only baseline, and 5.6% over standard conditional decoding of the same CFG-trained model. Oracle analysis shows that substantially larger WER reductions are possible through utterance-level scale selection and identifies how beneficial adjustments vary with domain shifts.
Asymmetric Classifier-Free Guidance for Target-Speaker ASR
Researchers introduced asymmetric classifier-free guidance (CFG) for target-speaker automatic speech recognition using Whisper, reporting relative word error rate reductions of up to 21.8% over the condition-only baseline and 5.6% over standard conditional decoding of the same CFG-trained model under domain shifts. The method uses a speaker-conditioned branch to predict the target transcript and a speaker-unconditioned branch to predict serialized multi-speaker transcripts, with a global guidance scale selected on target-domain development data and a lightweight encoder-based predictor adjusting it per utterance while the recognition model stays fixed. Oracle analysis indicates substantially larger WER reductions are possible through utterance-level scale selection.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.