{"slug": "learning-when-to-commit-from-partial-speech-for-end-to-end-simultaneous-speech", "title": "Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation", "summary": "Researchers adapted a full-utterance speech language model for simultaneous speech translation using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations, according to an arXiv paper (2610.02612v1). On FLEURS and CoVoST2 across three language directions, prefix training improved quality-latency frontiers over the unadapted model, and under multi-turn training commit-calibration error fell by 63-68% overall and 68-80% at early prefixes, while single-turn training gave only modest overall calibration gains and no early-prefix improvement. A confidence threshold controlled the inference-time quality-latency trade-off, with multi-turn append-only decoding generally stronger at low latency and a small synthesis margin sometimes extending the frontier to lower latency on shorter utterances.", "body_md": "arXiv:2610.02612v1 Announce Type: new \nAbstract: Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations. We compare single-turn forced-prefix and multi-turn append-only decoding, use a confidence threshold to control the inference-time quality--latency trade-off, and vary the density of training prefixes with a separate synthesis margin. On FLEURS and CoVoST2 in three language directions, prefix training improves quality--latency frontiers over the unadapted model, and confidence provides the broadest consistently competitive operating range. Multi-turn decoding is generally stronger at low latency; under multi-turn training, commit-calibration error falls by 63--68% overall and 68--80% at early prefixes, whereas single-turn training provides only modest overall calibration gains and no early-prefix improvement. A small synthesis margin sometimes extends the frontier to lower latency, particularly on shorter utterances, while a larger margin degrades translation quality and calibration. Prefix adaptation therefore improves simultaneous speech translation, especially under multi-turn append-only decoding, while synthesis density introduces a non-monotonic quality--latency trade-off.", "url": "https://wpnews.pro/news/learning-when-to-commit-from-partial-speech-for-end-to-end-simultaneous-speech", "canonical_source": "https://arxiv.org/abs/2610.02612", "published_at": "2026-10-05 04:00:00+00:00", "updated_at": "2026-10-05 04:15:01.127922+00:00", "lang": "en", "topics": ["natural-language-processing", "machine-learning", "artificial-intelligence", "ai-research"], "entities": ["arXiv", "FLEURS", "CoVoST2"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/learning-when-to-commit-from-partial-speech-for-end-to-end-simultaneous-speech", "markdown": "https://wpnews.pro/news/learning-when-to-commit-from-partial-speech-for-end-to-end-simultaneous-speech.md", "text": "https://wpnews.pro/news/learning-when-to-commit-from-partial-speech-for-end-to-end-simultaneous-speech.txt", "jsonld": "https://wpnews.pro/news/learning-when-to-commit-from-partial-speech-for-end-to-end-simultaneous-speech.jsonld"}}