Benchmarking Large Language Models for Game Localization Quality Assurance: A Cross-Model, Cross-Lingual Analysis A study presented at the 17th Conference of the Association for Machine Translation in the Americas (AMTA 2026) benchmarked eight large language models on game localization quality assurance tasks, finding that Claude Sonnet 4 achieved the best overall F1 score of 0.766, followed by Qwen-2.5-72B (0.711) and Gemini 2.0 Flash (0.691). The benchmark, spanning 96 evaluation settings and 48,000 translation samples across six target languages and two game genres, found that target language did not significantly affect performance (p = 0.285), French was the most consistent language, Japanese the most challenging, and game genre had minimal impact. The authors, Mao Tian and Na Wu, noted that open-weight Qwen-2.5-72B offers competitive quality at lower cost, providing guidance for deploying LLM-based LQA in production workflows. Abstract Localization quality assurance LQA is a critical component of game development, where manual review of large volumes of translated text is time-consuming and costly. Recent advances in large language models LLMs suggest strong potential for automated LQA, yet their effectiveness across different models, target languages, and game domains remains insufficiently understood. We present a comprehensive benchmark evaluating eight LLMs, including both closed-source and open-weight models, on English-to-six-language gaming LQA tasks across two game genres. Our dataset comprises 96 evaluation settings with a total of 48,000 translation samples. The results show that Claude Sonnet 4 achieves the best overall performance F1 = 0.766 , followed by Qwen-2.5-72B F1 = 0.711 and Gemini 2.0 Flash F1 = 0.691 . We observe that 1 the target language does not significantly affect model performance p = 0.285 , 2 models achieve their most consistent performance on French, while Japanese is the most challenging target language, and 3 game genre RPG vs. strategy has minimal impact on accuracy. While closed-source models achieve the highest overall performance, open-weight alternatives such as Qwen-2.5-72B provide competitive quality at substantially lower cost. These findings provide practical guidance for deploying LLM-based LQA systems in production game localization workflows.- Anthology ID: - 2026.amta-research.12 - Volume: Proceedings of the 17th Conference of the Association for Machine Translation in the Americas Volume 1: Research Track /volumes/2026.amta-research/ - Month: - August - Year: - 2026 - Address: - Québec City, Canada - Editors: Eleftheria Briakou /people/eleftheria-briakou/unverified/ , Jeremy Gwinnup /people/jeremy-gwinnup/ , Shivali Goel /people/shivali-goel/unverified/ - Venue: AMTA /venues/amta/ - SIG: - Publisher: - Association for Machine Translation in the Americas - Note: - Pages: - 186–201 - Language: - URL: https://aclanthology.org/2026.amta-research.12/ https://aclanthology.org/2026.amta-research.12/ - DOI: - Cite ACL : - Mao Tian and Na Wu. 2026. Benchmarking Large Language Models for Game Localization Quality Assurance: A Cross-Model, Cross-Lingual Analysis https://aclanthology.org/2026.amta-research.12/ . In Proceedings of the 17th Conference of the Association for Machine Translation in the Americas Volume 1: Research Track , pages 186–201, Québec City, Canada. Association for Machine Translation in the Americas. - Cite Informal : Benchmarking Large Language Models for Game Localization Quality Assurance: A Cross-Model, Cross-Lingual Analysis https://aclanthology.org/2026.amta-research.12/ Tian & Wu, AMTA 2026 - PDF: https://aclanthology.org/2026.amta-research.12.pdf https://aclanthology.org/2026.amta-research.12.pdf