{"slug": "banglakontho-closing-the-long-form-gap-in-bangla-text-to-speech", "title": "BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech", "summary": "Researchers released BanglaKontho, a 20-hour single-speaker Bangla text-to-speech corpus of 7,050 segmented utterances at 24 kHz drawn from professional audiobook recordings, alongside a reusable Bangla text normalizer and full preprocessing pipeline. An MB-iSTFT-VITS baseline trained from scratch on the corpus reached 9.5% WER and 4.46 naturalness MOS, versus 16.0% WER and 3.16 MOS for the same architecture retrained on the 12-hour IndicTTS-Bn corpus. The corpus is released openly under CC BY-NC 4.0 to address the lack of long-form, consistent single-speaker Bangla speech data.", "body_md": "arXiv:2609.29146v1 Announce Type: new \nAbstract: Bangla, the seventh most spoken language in the world, remains under-resourced for neural text-to-speech. Public Bangla speech corpora are dominated by short read-prompt utterances collected for speech recognition, leaving long-form prosody and consistent single-speaker narration uncovered. We present BanglaKontho, a single-speaker Bangla TTS corpus of 20 hours derived from professional audiobook recordings: 7,050 segmented utterances with verified transcripts at 24 kHz. We also release a reusable Bangla text normalizer covering Bangladeshi-style digit grouping, currency and date expressions, Danda punctuation and Unicode normalization, together with the full preprocessing pipeline. An MB-iSTFT-VITS baseline trained from scratch reaches 9.5% WER and 4.46 naturalness MOS, against 16.0% and 3.16 for the same architecture retrained on the 12-hour IndicTTS-Bn corpus. The corpus is released openly under CC BY-NC 4.0.", "url": "https://wpnews.pro/news/banglakontho-closing-the-long-form-gap-in-bangla-text-to-speech", "canonical_source": "https://arxiv.org/abs/2609.29146", "published_at": "2026-09-25 04:00:00+00:00", "updated_at": "2026-09-25 04:01:01.716521+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "ai-research"], "entities": ["BanglaKontho", "IndicTTS-Bn", "MB-iSTFT-VITS", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/banglakontho-closing-the-long-form-gap-in-bangla-text-to-speech", "markdown": "https://wpnews.pro/news/banglakontho-closing-the-long-form-gap-in-bangla-text-to-speech.md", "text": "https://wpnews.pro/news/banglakontho-closing-the-long-form-gap-in-bangla-text-to-speech.txt", "jsonld": "https://wpnews.pro/news/banglakontho-closing-the-long-form-gap-in-bangla-text-to-speech.jsonld"}}