cd /news/artificial-intelligence/fd-vad-semantic-endpoint-detection-f… · home › topics › artificial-intelligence › article
[ARTICLE · art-142239] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

Researchers introduced FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions for full-duplex voice interaction, combining a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model. FD-VAD achieved the highest end-of-turn recall among qualifying systems on the TurnBench dev set at 0.853 with false positives at or below 0.10 in a zero-shot setting, outperforming strong streaming and non-streaming semantic turn classifiers across in-domain and conversational evaluations. The work shows semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.

by read1 min views1 publishedSep 30, 2026

arXiv:2609.35791v1 Announce Type: new Abstract: Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages. We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference. We further introduce confidence-gated endpoint commitment to control interruption versus delay and boundary-focused hard-negative sampling to improve decisions around ambiguous turn boundaries. Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on TurnBench dev set $0.853$ (at FP<=0.10) in a zero-shot setting. These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @fd-vad 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fd-vad-semantic-endp…] indexed:0 read:1min 2026-09-30 · —