cd /news/large-language-models/large-language-models-for-machine-tr… · home › topics › large-language-models › article
[ARTICLE · art-148021] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged

A new arXiv paper (2610.10918v1) evaluating large language models for machine translation quality annotation found that LLM agreement with human annotators exceeds inter-human agreement on some tasks but remains unreliable across most settings, varying substantially by annotation scheme, language pair and domain. The study compared LLM and human performance on Multidimensional Quality Metrics (MQM) and Error Span Annotation (ESA) across a long-context test set of 70 language pairs plus the public WMT23 and WMT25 data. Humans were most challenged by fine-grained MQM annotations and low-resource language pairs, while LLMs struggled with minor errors, wrong language variants and error span annotation, leading the authors to propose human-LLM collaborative annotation pipelines.

by read1 min views2 publishedOct 9, 2026

arXiv:2610.10918v1 Announce Type: new Abstract: Large Language Models (LLMs) are considered to be a more efficient and cost-effective alternative to human judgment for Machine Translation (MT) evaluation. With MT evaluation spanning a large number of language pairs, domains and levels of annotation granularity, LLMs must be thoroughly evaluated across these dimensions before being reliably used as alternatives to human evaluation. In this paper, we evaluate the performance of LLMs for two prominent MT quality evaluation schemes: Multidimensional Quality Metrics (MQM) and Error Span Annotation (ESA) by comparing their agreement with human annotators. We present results on a long-context test set of 70 language pairs and the publicly available WMT23 and WMT25 data, investigating both score and error span annotation agreement across a variety of language pairs and domains. Our results show that while LLM agreement with human annotators exceeds agreement between human annotators for some evaluation tasks, both vary substantially across annotation schemes, language pairs and domains and remain unreliable for most settings. Furthermore, we identify challenges facing both human and LLM annotators: humans are particularly challenged by fine-grained MQM annotations and low-resource language pairs, while LLMs struggle with minor errors, wrong language variants and error span annotation. Our results highlight the potential for improvement for both human and LLM annotation performance, possibly through human-LLM collaborative annotation pipelines that address the reliability issues identified in this work.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/large-language-model…] indexed:0 read:1min 2026-10-09 · —