{"slug": "inference-time-target-speaker-unlearning-in-llm-based-automatic-speech", "title": "Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition", "summary": "Researchers introduced target-speaker unlearning ASR (TSU-ASR), a fully end-to-end framework that lets a frozen dual-stream speech LLM transcribe multi-speaker audio while excluding opt-out speakers, using a light-weight Enrollment-Conditioned Gating (ECG) module that works at inference time for speakers unseen during ECG training. On the AMI (English) dataset, transcription accuracy for opt-out words fell from 72.3% to 48.2%, and on AliMeeting (Mandarin) opt-out character accuracy fell from 73.6% to 27.3%, while retained speakers' error rates stayed roughly the same. The authors position the approach as a privacy-preserving option for video conferencing platforms, letting speakers opt out of automated AI transcription without leaving the meeting.", "body_md": "arXiv:2609.30439v1 Announce Type: new \nAbstract: We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech transcribed, the task requires an ASR system to transcribe all speakers except the opt-out ones, while still indicating when those speakers are active. As a first step towards tackling this task, we introduce a novel, light-weight Enrollment-Conditioned Gating (ECG) module attachable to a frozen dual-stream speech LLM that enables ASR for new opt-out speakers dynamically during inference, even those who were not seen during initial ECG training phase. Our experiments on both AMI (English) and AliMeeting (Mandarin) datasets show that speech transcription accuracy for corresponding opt-out words or characters falls from 72.3% to 48.2% and from 73.6% to 27.3%, respectively, while retained speakers' transcription error rates maintain more or less the same. Our approach provides a practical solution for modern video conferencing platforms, allowing speakers to dynamically opt-out from automated AI transcriptions without forcefully leaving the meeting sessions, enabling a privacy-preserving interface for potentially millions of online meetings daily.", "url": "https://wpnews.pro/news/inference-time-target-speaker-unlearning-in-llm-based-automatic-speech", "canonical_source": "https://arxiv.org/abs/2609.30439", "published_at": "2026-09-28 04:00:00+00:00", "updated_at": "2026-09-28 04:18:38.019081+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "natural-language-processing", "ai-safety", "ai-ethics"], "entities": ["TSU-ASR", "Enrollment-Conditioned Gating", "AMI", "AliMeeting", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/inference-time-target-speaker-unlearning-in-llm-based-automatic-speech", "markdown": "https://wpnews.pro/news/inference-time-target-speaker-unlearning-in-llm-based-automatic-speech.md", "text": "https://wpnews.pro/news/inference-time-target-speaker-unlearning-in-llm-based-automatic-speech.txt", "jsonld": "https://wpnews.pro/news/inference-time-target-speaker-unlearning-in-llm-based-automatic-speech.jsonld"}}