cd /news/artificial-intelligence/asymmetric-classifier-free-guidance-… · home › topics › artificial-intelligence › article
[ARTICLE · art-140797] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Asymmetric Classifier-Free Guidance for Target-Speaker ASR

Researchers introduced asymmetric classifier-free guidance (CFG) for target-speaker automatic speech recognition using Whisper, reporting relative word error rate reductions of up to 21.8% over the condition-only baseline and 5.6% over standard conditional decoding of the same CFG-trained model under domain shifts. The method uses a speaker-conditioned branch to predict the target transcript and a speaker-unconditioned branch to predict serialized multi-speaker transcripts, with a global guidance scale selected on target-domain development data and a lightweight encoder-based predictor adjusting it per utterance while the recognition model stays fixed. Oracle analysis indicates substantially larger WER reductions are possible through utterance-level scale selection.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.30476v1 Announce Type: cross Abstract: Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized multi-speaker transcripts. CFG adjusts the contribution of speaker conditioning during decoding through a single guidance scale. We select a global guidance scale on target-domain development data and train a lightweight encoder-based predictor to adjust it for each utterance, keeping the recognition model fixed. Under domain shifts, our full system achieves relative word error rate (WER) reductions of up to 21.8% over the condition-only baseline, and 5.6% over standard conditional decoding of the same CFG-trained model. Oracle analysis shows that substantially larger WER reductions are possible through utterance-level scale selection and identifies how beneficial adjustments vary with domain shifts.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @whisper 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/asymmetric-classifie…] indexed:0 read:1min 2026-09-28 · —