cd /news/computer-vision/curriculum-based-noise-adaptation-fo… · home topics computer-vision article
[ARTICLE · art-135519] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition

A new arXiv paper (2609.20839v1) proposes progressive error curriculum training (PECT), a curriculum learning framework that adapts a No Language Left Behind (NLLB)-based phoneme-to-text reconstruction model using synthetic phoneme perturbations and multi-domain and target-domain pseudo-labels generated by a visual speech recognizer. On the LRS2 and LRS3 benchmarks, PECT reduced the word error rate of HP-VSR-FiLMFuse (L4) from 23.3% to 22.2% on LRS2 and of HP-VSR-ResFiLM from 30.3% to 29.7% on LRS3, improving reconstruction across V-ASR, PV-ASR, and HP-VSR frontends.

by read1 min views1 publishedSep 21, 2026

arXiv:2609.20839v1 Announce Type: new Abstract: Phoneme-centric visual speech recognition reconstructs sentences from intermediate phoneme predictions, making overall recognition performance highly dependent on the robustness of the phoneme-to-text reconstruction model. Existing reconstruction approaches are commonly trained on clean phoneme sequences or synthetically corrupted inputs, leading to a mismatch between training conditions and the realistic phoneme prediction errors encountered during inference. To address this limitation, this paper proposes progressive error curriculum training (PECT). This curriculum learning framework progressively adapts a No Language Left Behind (NLLB)-based phoneme-to-text reconstruction model using synthetic phoneme perturbations, multi-domain pseudo-labels, and target-domain pseudo-labels generated by a visual speech recognizer. By gradually exposing the reconstruction model to increasingly realistic phoneme prediction errors, the proposed framework improves robustness while preserving sentence-reconstruction accuracy. Experiments on the LRS2 and LRS3 benchmarks demonstrate that PECT consistently improves reconstruction performance across multiple phoneme-based visual speech recognition frontends, including visual automatic speech recognition (V-ASR), point visual automatic speech recognition (PV-ASR), and head-pose-aware visual speech recognition (HP-VSR) variants. In particular, PECT reduces the word error rate (WER) of HP-VSR-FiLMFuse (L4) from 23.3% to 22.2% on LRS2 and reduces the WER of HP-VSR-ResFiLM from 30.3% to 29.7% on LRS3. Comprehensive ablation studies and qualitative analyses further demonstrate the effectiveness of progressively adapting the reconstruction model to realistic phoneme prediction errors. These results show that PECT provides an effective and generalizable curriculum learning strategy for phoneme-to-text reconstruction in phoneme-centric visual speech recognition.

── more in #computer-vision 4 stories · sorted by recency
── more on @no language left behind (nllb) 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/curriculum-based-noi…] indexed:0 read:1min 2026-09-21 ·