FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech Researchers introduced FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions for full-duplex voice interaction, combining a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model. FD-VAD achieved the highest end-of-turn recall among qualifying systems on the TurnBench dev set at 0.853 with false positives at or below 0.10 in a zero-shot setting, outperforming strong streaming and non-streaming semantic turn classifiers across in-domain and conversational evaluations. The work shows semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking. arXiv:2609.35791v1 Announce Type: new Abstract: Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages. We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference. We further introduce confidence-gated endpoint commitment to control interruption versus delay and boundary-focused hard-negative sampling to improve decisions around ambiguous turn boundaries. Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on TurnBench dev set $0.853$ at FP<=0.10 in a zero-shot setting. These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.