{"slug": "a-systematic-evaluation-of-cross-lingual-consistency-enhancement-methods-in", "title": "A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models", "summary": "A new arXiv study (2609.04409v1) systematically evaluated cross-lingual consistency (CLC) enhancement methods in multilingual language models, finding that post-training methods, especially direct distribution alignment, consistently improve CLC across three model families and three closed-form benchmarks, while other methods are sensitive to answer format and language coverage. The study also found no systematic degradation in culturally diverse question answering under controlled closed-form evaluation, though open-ended generation showed occasional accuracy reductions for non-English responses.", "body_md": "arXiv:2609.04409v1 Announce Type: new \nAbstract: Multilingual language models often produce inconsistent answers to semantically equivalent questions across languages, motivating methods to improve cross-lingual consistency (CLC). However, existing methods are typically evaluated using different models, tasks, and protocols, leaving their relative strengths unclear. In this work, we present a unified evaluation of representative CLC-enhancement methods for question answering, spanning inference-time interventions and post-training approaches across three model families and three closed-form benchmarks. The results show that post-training methods are generally more reliable, with direct distribution alignment consistently improving CLC across all model-dataset combinations, while other methods are more sensitive to answer format and the breadth of language coverage. Notably, cross-domain transfer is limited unless source and target tasks share similar output formats. We further investigate whether CLC enhancement hurts models' ability to respond differently *when needed*, that is, when asked culture-dependent questions. Across two benchmarks of culturally diverse question answering, we find no systematic degradation in controlled closed-form evaluation, whereas open-ended generation reveals occasional accuracy reductions, particularly for non-English responses. Our work highlights the need to evaluate CLC enhancement for both cross-domain robustness and culturally appropriate variation, informing future work in post-training and benchmark development.", "url": "https://wpnews.pro/news/a-systematic-evaluation-of-cross-lingual-consistency-enhancement-methods-in", "canonical_source": "https://arxiv.org/abs/2609.04409", "published_at": "2026-09-07 04:00:00+00:00", "updated_at": "2026-09-07 04:25:33.690620+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/a-systematic-evaluation-of-cross-lingual-consistency-enhancement-methods-in", "markdown": "https://wpnews.pro/news/a-systematic-evaluation-of-cross-lingual-consistency-enhancement-methods-in.md", "text": "https://wpnews.pro/news/a-systematic-evaluation-of-cross-lingual-consistency-enhancement-methods-in.txt", "jsonld": "https://wpnews.pro/news/a-systematic-evaluation-of-cross-lingual-consistency-enhancement-methods-in.jsonld"}}