arXiv:2609.27205v1 Announce Type: new Abstract: Text-to-speech systems increasingly process user-generated text (UGT) such as ppl and imo, whose pronunciation must be inferred from the canonical rather than surface form. We introduce UGTPhon, the first grapheme-to-phoneme (G2P) benchmark for UGT in English, Vietnamese, and Korean, together with an inference-grounded taxonomy for fine-grained diagnosis. Existing G2P models and frontier LLMs exhibit a systematic canonical-to-non-canonical performance gap, reaching up to 66.8 PER points. As a benchmark baseline, we propose a simple compositional G2P approach that incorporates canonical-form evidence through exact-match lookup and staged decoding. Across matched ByT5 and Qwen2.5-0.5B backbones, explicit canonical-form modeling consistently reduces non-canonical G2P errors. The 0.5B variant also performs competitively with much larger few-shot frontier LLMs, highlighting the benefit of explicitly modeling canonical-form inference for UGT phonemization.
Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach
Researchers introduced UGTPhon, described as the first grapheme-to-phoneme (G2P) benchmark for user-generated text in English, Vietnamese, and Korean, alongside an inference-grounded taxonomy for fine-grained diagnosis. Existing G2P models and frontier LLMs show a systematic canonical-to-non-canonical performance gap of up to 66.8 PER points, while a compositional G2P baseline using exact-match lookup and staged decoding reduced non-canonical errors across matched ByT5 and Qwen2.5-0.5B backbones, with the 0.5B variant performing competitively with much larger few-shot frontier LLMs.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.