cd /news/large-language-models/benchmarking-large-language-models-f… · home topics large-language-models article
[ARTICLE · art-118862] src=aclanthology.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Benchmarking Large Language Models for Game Localization Quality Assurance: A Cross-Model, Cross-Lingual Analysis

A study presented at the 17th Conference of the Association for Machine Translation in the Americas (AMTA 2026) benchmarked eight large language models on game localization quality assurance tasks, finding that Claude Sonnet 4 achieved the best overall F1 score of 0.766, followed by Qwen-2.5-72B (0.711) and Gemini 2.0 Flash (0.691). The benchmark, spanning 96 evaluation settings and 48,000 translation samples across six target languages and two game genres, found that target language did not significantly affect performance (p = 0.285), French was the most consistent language, Japanese the most challenging, and game genre had minimal impact. The authors, Mao Tian and Na Wu, noted that open-weight Qwen-2.5-72B offers competitive quality at lower cost, providing guidance for deploying LLM-based LQA in production workflows.

read2 min views11 publishedSep 1, 2026
Benchmarking Large Language Models for Game Localization Quality Assurance: A Cross-Model, Cross-Lingual Analysis
Image: Aclanthology (auto-discovered)
Abstract

Localization quality assurance (LQA) is a critical component of game development, where manual review of large volumes of translated text is time-consuming and costly. Recent advances in large language models (LLMs) suggest strong potential for automated LQA, yet their effectiveness across different models, target languages, and game domains remains insufficiently understood. We present a comprehensive benchmark evaluating eight LLMs, including both closed-source and open-weight models, on English-to-six-language gaming LQA tasks across two game genres. Our dataset comprises 96 evaluation settings with a total of 48,000 translation samples. The results show that Claude Sonnet 4 achieves the best overall performance (F1 = 0.766), followed by Qwen-2.5-72B (F1 = 0.711) and Gemini 2.0 Flash (F1 = 0.691). We observe that (1) the target language does not significantly affect model performance (p = 0.285), (2) models achieve their most consistent performance on French, while Japanese is the most challenging target language, and (3) game genre (RPG vs. strategy) has minimal impact on accuracy. While closed-source models achieve the highest overall performance, open-weight alternatives such as Qwen-2.5-72B provide competitive quality at substantially lower cost. These findings provide practical guidance for deploying LLM-based LQA systems in production game localization workflows.- Anthology ID:

- 2026.amta-research.12
- Volume:
[Proceedings of the 17th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track)](/volumes/2026.amta-research/)- Month:
  • August
  • Year:
  • 2026
  • Address:
  • Québec City, Canada
- Editors:
[Eleftheria Briakou](/people/eleftheria-briakou/unverified/),[Jeremy Gwinnup](/people/jeremy-gwinnup/),[Shivali Goel](/people/shivali-goel/unverified/)- Venue:
[AMTA](/venues/amta/)- SIG:
- Publisher:
  • Association for Machine Translation in the Americas
- Note:
- Pages:
  • 186–201
- Language:
- URL:
[https://aclanthology.org/2026.amta-research.12/](https://aclanthology.org/2026.amta-research.12/)- DOI:
- Cite (ACL):
[Benchmarking Large Language Models for Game Localization Quality Assurance: A Cross-Model, Cross-Lingual Analysis](https://aclanthology.org/2026.amta-research.12/)(Tian & Wu, AMTA 2026)- PDF:
[https://aclanthology.org/2026.amta-research.12.pdf](https://aclanthology.org/2026.amta-research.12.pdf)
── more in #large-language-models 4 stories · sorted by recency
── more on @claude sonnet 4 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/benchmarking-large-l…] indexed:0 read:2min 2026-09-01 ·