cd /news/artificial-intelligence/building-and-evaluating-a-synthetic-… · home topics artificial-intelligence article
[ARTICLE · art-108245] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

Researchers released a synthetic Bengali speech dataset for telecom customer-care scenarios, containing 10,000 audio-text pairs (26.82 hours of 24 kHz speech) with predefined splits of 9,000/500/500 for train, validation, and test. Generated with OmniVoice in voice-cloning mode using a real female reference recording, the dataset is publicly available on Hugging Face under CC-BY-4.0 and includes normalized transcripts for ASR/STT training. Evaluation using a domain-adapted Whisper ASR model yielded an average WER of 2.54% and CER of 0.59%, with median values of 0.00%, indicating strong text-audio consistency.

read1 min views3 publishedAug 24, 2026

arXiv:2608.20346v1 Announce Type: new Abstract: Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs, approximately 26.82 hours of 24 kHz speech, and predefined train, validation, and test splits of 9,000, 500, and 500 examples. It is publicly released on Hugging Face under the CC-BY-4.0 license. The speech was generated with OmniVoice in voice-cloning mode using a real female reference recording and transcript, with bfloat16 precision, 16 diffusion sampling steps, and a speaking-rate control value of 1.0. Along with the original Bengali text, the dataset provides a normalized transcript field designed for ASR/STT training and evaluation. We report an automatic intelligibility check over all 10,000 samples using a domain-adapted Whisper ASR model fine-tuned from bengaliAI/tugstugi_bengaliai-regional-asr_whisper-medium, along with a manual listening check on selected samples. The evaluation gives an average WER of 2.54%, an average CER of 0.59%, and median WER and CER values of 0.00%. These results suggest strong text-audio consistency under the selected automatic evaluation pipeline, while the paper also discusses the limitations of synthetic speech and STT-based evaluation.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @omnivoice 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/building-and-evaluat…] indexed:0 read:1min 2026-08-24 ·