{"slug": "phonemizing-user-generated-text-a-benchmark-taxonomy-and-compositional-approach", "title": "Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach", "summary": "Researchers introduced UGTPhon, described as the first grapheme-to-phoneme (G2P) benchmark for user-generated text in English, Vietnamese, and Korean, alongside an inference-grounded taxonomy for fine-grained diagnosis. Existing G2P models and frontier LLMs show a systematic canonical-to-non-canonical performance gap of up to 66.8 PER points, while a compositional G2P baseline using exact-match lookup and staged decoding reduced non-canonical errors across matched ByT5 and Qwen2.5-0.5B backbones, with the 0.5B variant performing competitively with much larger few-shot frontier LLMs.", "body_md": "arXiv:2609.27205v1 Announce Type: new \nAbstract: Text-to-speech systems increasingly process user-generated text (UGT) such as ppl and imo, whose pronunciation must be inferred from the canonical rather than surface form. We introduce UGTPhon, the first grapheme-to-phoneme (G2P) benchmark for UGT in English, Vietnamese, and Korean, together with an inference-grounded taxonomy for fine-grained diagnosis. Existing G2P models and frontier LLMs exhibit a systematic canonical-to-non-canonical performance gap, reaching up to 66.8 PER points. As a benchmark baseline, we propose a simple compositional G2P approach that incorporates canonical-form evidence through exact-match lookup and staged decoding. Across matched ByT5 and Qwen2.5-0.5B backbones, explicit canonical-form modeling consistently reduces non-canonical G2P errors. The 0.5B variant also performs competitively with much larger few-shot frontier LLMs, highlighting the benefit of explicitly modeling canonical-form inference for UGT phonemization.", "url": "https://wpnews.pro/news/phonemizing-user-generated-text-a-benchmark-taxonomy-and-compositional-approach", "canonical_source": "https://arxiv.org/abs/2609.27205", "published_at": "2026-09-24 04:00:00+00:00", "updated_at": "2026-09-24 04:32:59.706967+00:00", "lang": "en", "topics": ["natural-language-processing", "machine-learning", "ai-research", "large-language-models"], "entities": ["UGTPhon", "ByT5", "Qwen2.5-0.5B", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/phonemizing-user-generated-text-a-benchmark-taxonomy-and-compositional-approach", "markdown": "https://wpnews.pro/news/phonemizing-user-generated-text-a-benchmark-taxonomy-and-compositional-approach.md", "text": "https://wpnews.pro/news/phonemizing-user-generated-text-a-benchmark-taxonomy-and-compositional-approach.txt", "jsonld": "https://wpnews.pro/news/phonemizing-user-generated-text-a-benchmark-taxonomy-and-compositional-approach.jsonld"}}