cd /news/natural-language-processing/learning-when-to-commit-from-partial… · home › topics › natural-language-processing › article
[ARTICLE · art-145171] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

Researchers adapted a full-utterance speech language model for simultaneous speech translation using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations, according to an arXiv paper (2610.02612v1). On FLEURS and CoVoST2 across three language directions, prefix training improved quality-latency frontiers over the unadapted model, and under multi-turn training commit-calibration error fell by 63-68% overall and 68-80% at early prefixes, while single-turn training gave only modest overall calibration gains and no early-prefix improvement. A confidence threshold controlled the inference-time quality-latency trade-off, with multi-turn append-only decoding generally stronger at low latency and a small synthesis margin sometimes extending the frontier to lower latency on shorter utterances.

by read1 min views1 publishedOct 5, 2026

arXiv:2610.02612v1 Announce Type: new Abstract: Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations. We compare single-turn forced-prefix and multi-turn append-only decoding, use a confidence threshold to control the inference-time quality--latency trade-off, and vary the density of training prefixes with a separate synthesis margin. On FLEURS and CoVoST2 in three language directions, prefix training improves quality--latency frontiers over the unadapted model, and confidence provides the broadest consistently competitive operating range. Multi-turn decoding is generally stronger at low latency; under multi-turn training, commit-calibration error falls by 63--68% overall and 68--80% at early prefixes, whereas single-turn training provides only modest overall calibration gains and no early-prefix improvement. A small synthesis margin sometimes extends the frontier to lower latency, particularly on shorter utterances, while a larger margin degrades translation quality and calibration. Prefix adaptation therefore improves simultaneous speech translation, especially under multi-turn append-only decoding, while synthesis density introduces a non-monotonic quality--latency trade-off.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/learning-when-to-com…] indexed:0 read:1min 2026-10-05 · —