cd /news/artificial-intelligence/averaging-bias-human-faithfulness-an… · home topics artificial-intelligence article
[ARTICLE · art-85567] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Averaging Bias: Human Faithfulness Annotations are not Locally Faithful

A new study from arXiv (2608.00205v1) finds that human faithfulness annotations for text summarization exhibit an 'Averaging Bias,' where annotators accept summaries as faithful if most sentences are supported, rather than requiring all sentences to be supported. Using five large language model (LLM) judges across four benchmarks, the authors show that global human labels correlate better with the average of per-sentence LLM judgments than with a strict conjunctive rule, and manual review confirms that many human-labeled faithful summaries contain genuine local factual errors.

read1 min views1 publishedAug 4, 2026

arXiv:2608.00205v1 Announce Type: new Abstract: Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if every of its sentences is supported by the source document: a strict conjunctive rule under which a single unsupported sentence makes the whole summary unfaithful. Yet most faithfulness benchmarks collect only one global human annotation label per summary. We ask whether such global human labels actually implement the conjunctive rule. We hypothesize that annotators may accept a summary as faithful when most sentences are faithful, not only when all are faithful. To test our hypothesis, we use five large language model (LLM) judges as per-sentence raters across four widely used faithfulness benchmarks. We find that global human labels correlate better with the average of per-sentence LLM judgments than with the implementation of the strict conjunctive rule. A manual review confirms that a substantial fraction of summaries labeled faithful by humans contain genuine local factual errors. We call this tendency Averaging Bias. Our results reveal that human labels on widely used faithfulness benchmarks contain measurable Averaging Bias, calling for carefully structured designs for trustworthy human annotations

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/averaging-bias-human…] indexed:0 read:1min 2026-08-04 ·