cd /news/large-language-models/evaluating-the-evaluator-summarizati… · home topics large-language-models article
[ARTICLE · art-119833] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English

A new multilingual summary meta-evaluation dataset (BASSE) containing human judgments on 2,040 abstractive summaries reveals that proprietary judge LLMs correlate most strongly with human ratings, followed by criteria-specific automatic metrics, while open-sourced judge LLMs perform poorly. The dataset, described in arXiv paper 2503.17039v3, includes summaries generated manually or by five large language models with four prompts, each evaluated on coherence, consistency, fluency, relevance, and 5W1H using a 5-point Likert scale.

read1 min views3 publishedSep 3, 2026

arXiv:2503.17039v3 Announce Type: replace-cross Abstract: Automatic text summarization relies on automatic evaluation to quickly determine the quality of summarization models via automatic metrics and LLM-as-a-Judge models. However, these techniques require meta-evaluation to ensure that they capture human judgments correctly. In this paper, we explore this meta-evaluation beyond English by generating a new multilingual summary meta-evaluation dataset (BASSE), which comprises human judgments on 2,040 abstractive summaries, generated either manually or by five Large Language Models (LLMs) with four different prompts. For each summary, annotators evaluate five criteria on a 5-point Likert scale: coherence, consistency, fluency, relevance, and 5W1H. We then benchmark automatic summarization metrics and LLM-as-a-Judge models. Our results show that currently proprietary judge LLMs have the highest correlation with human judgments, followed by criteria-specific automatic metrics, while open-sourced judge LLMs perform poorly.

── more in #large-language-models 4 stories · sorted by recency
── more on @basse 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evaluating-the-evalu…] indexed:0 read:1min 2026-09-03 ·