{"slug": "evaluating-the-evaluator-summarization-metrics-and-llm-judges-beyond-english", "title": "Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English", "summary": "A new multilingual summary meta-evaluation dataset (BASSE) containing human judgments on 2,040 abstractive summaries reveals that proprietary judge LLMs correlate most strongly with human ratings, followed by criteria-specific automatic metrics, while open-sourced judge LLMs perform poorly. The dataset, described in arXiv paper 2503.17039v3, includes summaries generated manually or by five large language models with four prompts, each evaluated on coherence, consistency, fluency, relevance, and 5W1H using a 5-point Likert scale.", "body_md": "arXiv:2503.17039v3 Announce Type: replace-cross\nAbstract: Automatic text summarization relies on automatic evaluation to quickly determine the quality of summarization models via automatic metrics and LLM-as-a-Judge models. However, these techniques require meta-evaluation to ensure that they capture human judgments correctly. In this paper, we explore this meta-evaluation beyond English by generating a new multilingual summary meta-evaluation dataset (BASSE), which comprises human judgments on 2,040 abstractive summaries, generated either manually or by five Large Language Models (LLMs) with four different prompts. For each summary, annotators evaluate five criteria on a 5-point Likert scale: coherence, consistency, fluency, relevance, and 5W1H. We then benchmark automatic summarization metrics and LLM-as-a-Judge models. Our results show that currently proprietary judge LLMs have the highest correlation with human judgments, followed by criteria-specific automatic metrics, while open-sourced judge LLMs perform poorly.", "url": "https://wpnews.pro/news/evaluating-the-evaluator-summarization-metrics-and-llm-judges-beyond-english", "canonical_source": "https://www.machinebrief.com/news/evaluating-the-evaluator-summarization-metrics-and-llm-judge-3qrc", "published_at": "2026-09-03 04:00:00+00:00", "updated_at": "2026-09-03 04:53:06.940054+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "ai-research"], "entities": ["BASSE", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/evaluating-the-evaluator-summarization-metrics-and-llm-judges-beyond-english", "markdown": "https://wpnews.pro/news/evaluating-the-evaluator-summarization-metrics-and-llm-judges-beyond-english.md", "text": "https://wpnews.pro/news/evaluating-the-evaluator-summarization-metrics-and-llm-judges-beyond-english.txt", "jsonld": "https://wpnews.pro/news/evaluating-the-evaluator-summarization-metrics-and-llm-judges-beyond-english.jsonld"}}