{"slug": "a-factorial-study-of-synthetic-data-generation-for-low-resource-machine-using", "title": "A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books", "summary": "A new pipeline using large language models to extract grammatical rules and lexicons from grammar books generates synthetic parallel corpora for fine-tuning machine translation models, improving over seed-data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, with best-case ChrF++ gains of +8.8, +5.3, and +3.3 respectively, according to a factorial study on three typologically diverse low-resource languages.", "body_md": "arXiv:2607.22376v1 Announce Type: new\nAbstract: Most endangered languages lack the parallel data required for machine translation, despite the existence of descriptive grammar books. We introduce a pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic parallel corpora for fine-tuning-rather than feeding grammar content into prompts at inference time, as in prior work. Validated on three typologically diverse low-resource languages-Kalamang (Papuan), Tuatschin (Romance), and Mandan (Siouan)-we show that fine-tuning on synthetic data improves over seed-data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, with best-case ChrF++ gains of +8.8, +5.3, and +3.3 respectively. Through a systematic factorial study across 96 configurations varying target part-of-speech, retrieval granularity, and sample volume, we identify which factor combinations drive gains and where they break down. Our results demonstrate that static linguistic documentation can be repurposed for machine translation fine-tuning, offering a practical path towards translation tools for severely under-resourced languages.", "url": "https://wpnews.pro/news/a-factorial-study-of-synthetic-data-generation-for-low-resource-machine-using", "canonical_source": "https://www.machinebrief.com/news/a-factorial-study-of-synthetic-data-generation-for-low-resou-86qt", "published_at": "2026-07-27 04:00:00+00:00", "updated_at": "2026-07-27 05:27:59.042759+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "large-language-models"], "entities": ["Kalamang", "Tuatschin", "Mandan"], "alternates": {"html": "https://wpnews.pro/news/a-factorial-study-of-synthetic-data-generation-for-low-resource-machine-using", "markdown": "https://wpnews.pro/news/a-factorial-study-of-synthetic-data-generation-for-low-resource-machine-using.md", "text": "https://wpnews.pro/news/a-factorial-study-of-synthetic-data-generation-for-low-resource-machine-using.txt", "jsonld": "https://wpnews.pro/news/a-factorial-study-of-synthetic-data-generation-for-low-resource-machine-using.jsonld"}}