{"slug": "fd-vad-semantic-endpoint-detection-for-streaming-full-duplex-speech", "title": "FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech", "summary": "Researchers introduced FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions for full-duplex voice interaction, combining a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model. FD-VAD achieved the highest end-of-turn recall among qualifying systems on the TurnBench dev set at 0.853 with false positives at or below 0.10 in a zero-shot setting, outperforming strong streaming and non-streaming semantic turn classifiers across in-domain and conversational evaluations. The work shows semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.", "body_md": "arXiv:2609.35791v1 Announce Type: new \nAbstract: Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages. We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference. We further introduce confidence-gated endpoint commitment to control interruption versus delay and boundary-focused hard-negative sampling to improve decisions around ambiguous turn boundaries. Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on TurnBench dev set $0.853$ (at FP<=0.10) in a zero-shot setting. These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.", "url": "https://wpnews.pro/news/fd-vad-semantic-endpoint-detection-for-streaming-full-duplex-speech", "canonical_source": "https://arxiv.org/abs/2609.35791", "published_at": "2026-09-30 04:00:00+00:00", "updated_at": "2026-09-30 04:19:58.468549+00:00", "lang": "en", "topics": ["artificial-intelligence", "natural-language-processing", "machine-learning", "ai-research"], "entities": ["FD-VAD", "TurnBench", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/fd-vad-semantic-endpoint-detection-for-streaming-full-duplex-speech", "markdown": "https://wpnews.pro/news/fd-vad-semantic-endpoint-detection-for-streaming-full-duplex-speech.md", "text": "https://wpnews.pro/news/fd-vad-semantic-endpoint-detection-for-streaming-full-duplex-speech.txt", "jsonld": "https://wpnews.pro/news/fd-vad-semantic-endpoint-detection-for-streaming-full-duplex-speech.jsonld"}}