arXiv:2610.09321v1 Announce Type: new Abstract: Speech Language Model (SLM) performance often degrades on dialects due to data scarcity. Conventional text-to-speech (TTS) augmentation struggles to cover diverse dialects as it requires a certain amount of real dialect speech. We propose synthesizing pseudo-dialect speech by converting LLM-generated dialect text via a standard-language TTS model, requiring zero real dialect speech. Additionally, we introduce intermediate standard-text prediction during training, acting as semantic normalization for downstream tasks. We evaluate dialect understanding via dialect-to-English speech translation across Japanese, German, and Chinese dialects. Compared to synthetic standard speech baselines, pseudo-dialect augmentation improves scores for Japanese (from 25.38 to 26.24) and German (from 31.57 to 32.47). Furthermore, the intermediate standard-text prediction effectively bridges the semantic gap, boosting performance to 28.26 for Japanese and from 11.67 to 16.37 for Chinese. These results suggest that our approach scales to various languages without requiring speech resources specific to each dialect.
Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation
Synthesizing pseudo-dialect speech from LLM-generated dialect text passed through a standard-language TTS model improved dialect-to-English speech translation scores without any real dialect speech, according to an arXiv paper (2610.09321v1). Pseudo-dialect augmentation raised Japanese scores from 25.38 to 26.24 and German from 31.57 to 32.47 over synthetic standard-speech baselines, while adding intermediate standard-text prediction during training lifted Japanese to 28.26 and Chinese from 11.67 to 16.37. The authors report the approach scales across Japanese, German, and Chinese dialects without requiring dialect-specific speech resources.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.