BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech Researchers released BanglaKontho, a 20-hour single-speaker Bangla text-to-speech corpus of 7,050 segmented utterances at 24 kHz drawn from professional audiobook recordings, alongside a reusable Bangla text normalizer and full preprocessing pipeline. An MB-iSTFT-VITS baseline trained from scratch on the corpus reached 9.5% WER and 4.46 naturalness MOS, versus 16.0% WER and 3.16 MOS for the same architecture retrained on the 12-hour IndicTTS-Bn corpus. The corpus is released openly under CC BY-NC 4.0 to address the lack of long-form, consistent single-speaker Bangla speech data. arXiv:2609.29146v1 Announce Type: new Abstract: Bangla, the seventh most spoken language in the world, remains under-resourced for neural text-to-speech. Public Bangla speech corpora are dominated by short read-prompt utterances collected for speech recognition, leaving long-form prosody and consistent single-speaker narration uncovered. We present BanglaKontho, a single-speaker Bangla TTS corpus of 20 hours derived from professional audiobook recordings: 7,050 segmented utterances with verified transcripts at 24 kHz. We also release a reusable Bangla text normalizer covering Bangladeshi-style digit grouping, currency and date expressions, Danda punctuation and Unicode normalization, together with the full preprocessing pipeline. An MB-iSTFT-VITS baseline trained from scratch reaches 9.5% WER and 4.46 naturalness MOS, against 16.0% and 3.16 for the same architecture retrained on the 12-hour IndicTTS-Bn corpus. The corpus is released openly under CC BY-NC 4.0.