{"slug": "curriculum-based-noise-adaptation-for-phoneme-to-text-reconstruction-in-visual", "title": "Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition", "summary": "A new arXiv paper (2609.20839v1) proposes progressive error curriculum training (PECT), a curriculum learning framework that adapts a No Language Left Behind (NLLB)-based phoneme-to-text reconstruction model using synthetic phoneme perturbations and multi-domain and target-domain pseudo-labels generated by a visual speech recognizer. On the LRS2 and LRS3 benchmarks, PECT reduced the word error rate of HP-VSR-FiLMFuse (L4) from 23.3% to 22.2% on LRS2 and of HP-VSR-ResFiLM from 30.3% to 29.7% on LRS3, improving reconstruction across V-ASR, PV-ASR, and HP-VSR frontends.", "body_md": "arXiv:2609.20839v1 Announce Type: new \nAbstract: Phoneme-centric visual speech recognition reconstructs sentences from intermediate phoneme predictions, making overall recognition performance highly dependent on the robustness of the phoneme-to-text reconstruction model. Existing reconstruction approaches are commonly trained on clean phoneme sequences or synthetically corrupted inputs, leading to a mismatch between training conditions and the realistic phoneme prediction errors encountered during inference. To address this limitation, this paper proposes progressive error curriculum training (PECT). This curriculum learning framework progressively adapts a No Language Left Behind (NLLB)-based phoneme-to-text reconstruction model using synthetic phoneme perturbations, multi-domain pseudo-labels, and target-domain pseudo-labels generated by a visual speech recognizer. By gradually exposing the reconstruction model to increasingly realistic phoneme prediction errors, the proposed framework improves robustness while preserving sentence-reconstruction accuracy. Experiments on the LRS2 and LRS3 benchmarks demonstrate that PECT consistently improves reconstruction performance across multiple phoneme-based visual speech recognition frontends, including visual automatic speech recognition (V-ASR), point visual automatic speech recognition (PV-ASR), and head-pose-aware visual speech recognition (HP-VSR) variants. In particular, PECT reduces the word error rate (WER) of HP-VSR-FiLMFuse (L4) from 23.3% to 22.2% on LRS2 and reduces the WER of HP-VSR-ResFiLM from 30.3% to 29.7% on LRS3. Comprehensive ablation studies and qualitative analyses further demonstrate the effectiveness of progressively adapting the reconstruction model to realistic phoneme prediction errors. These results show that PECT provides an effective and generalizable curriculum learning strategy for phoneme-to-text reconstruction in phoneme-centric visual speech recognition.", "url": "https://wpnews.pro/news/curriculum-based-noise-adaptation-for-phoneme-to-text-reconstruction-in-visual", "canonical_source": "https://arxiv.org/abs/2609.20839", "published_at": "2026-09-21 04:00:00+00:00", "updated_at": "2026-09-21 04:23:35.929429+00:00", "lang": "en", "topics": ["computer-vision", "natural-language-processing", "machine-learning", "ai-research"], "entities": ["No Language Left Behind (NLLB)", "LRS2", "LRS3", "V-ASR", "PV-ASR", "HP-VSR", "HP-VSR-FiLMFuse (L4)", "HP-VSR-ResFiLM"], "alternates": {"html": "https://wpnews.pro/news/curriculum-based-noise-adaptation-for-phoneme-to-text-reconstruction-in-visual", "markdown": "https://wpnews.pro/news/curriculum-based-noise-adaptation-for-phoneme-to-text-reconstruction-in-visual.md", "text": "https://wpnews.pro/news/curriculum-based-noise-adaptation-for-phoneme-to-text-reconstruction-in-visual.txt", "jsonld": "https://wpnews.pro/news/curriculum-based-noise-adaptation-for-phoneme-to-text-reconstruction-in-visual.jsonld"}}