Averaging Bias: Human Faithfulness Annotations are not Locally Faithful A new study from arXiv (2608.00205v1) finds that human faithfulness annotations for text summarization exhibit an 'Averaging Bias,' where annotators accept summaries as faithful if most sentences are supported, rather than requiring all sentences to be supported. Using five large language model (LLM) judges across four benchmarks, the authors show that global human labels correlate better with the average of per-sentence LLM judgments than with a strict conjunctive rule, and manual review confirms that many human-labeled faithful summaries contain genuine local factual errors. arXiv:2608.00205v1 Announce Type: new Abstract: Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if every of its sentences is supported by the source document: a strict conjunctive rule under which a single unsupported sentence makes the whole summary unfaithful. Yet most faithfulness benchmarks collect only one global human annotation label per summary. We ask whether such global human labels actually implement the conjunctive rule. We hypothesize that annotators may accept a summary as faithful when most sentences are faithful, not only when all are faithful. To test our hypothesis, we use five large language model LLM judges as per-sentence raters across four widely used faithfulness benchmarks. We find that global human labels correlate better with the average of per-sentence LLM judgments than with the implementation of the strict conjunctive rule. A manual review confirms that a substantial fraction of summaries labeled faithful by humans contain genuine local factual errors. We call this tendency Averaging Bias. Our results reveal that human labels on widely used faithfulness benchmarks contain measurable Averaging Bias, calling for carefully structured designs for trustworthy human annotations