cd /news/natural-language-processing/how-much-human-label-variation-does-… · home topics natural-language-processing article
[ARTICLE · art-65400] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI

A study measuring how much formal semantic structure explains human label variation in natural language inference found that it accounts for only 3.3 to 3.6 percent of entropy variance, with a weak median-split AUC of 0.606, indicating that formal semantics shifts disagreement amounts slightly but does not change what annotators disagree about. The analysis of 3,113 SNLI and MNLI items from ChaosNLI revealed a group-level boundary where non-purely upward monotone hypotheses show higher label entropy, but item-level ceilings and composition invariance were observed.

read1 min views1 publishedJul 20, 2026

arXiv:2607.15870v1 Announce Type: new Abstract: Human label variation in natural language inference is increasingly treated as signal rather than noise, but how much of it formal semantic structure explains has not been measured directly. We measure it on the 3,113 SNLI and MNLI items of ChaosNLI, using a rule-based operator and monotonicity tagger validated against MED (0.883 agreement at the edit site, 0.807 on the sentence-level summary our analyses consume), three preregistered analysis blocks, and full reporting of negative results. Three bounds emerge. First, a group-level boundary: hypotheses that are not purely upward monotone show reliably higher label entropy (Cliff's delta = -0.284), and rank-based tests defend the effect against operator-presence and length reductions, though a bounded-outcome sensitivity check weakens the regression form of the length defense. Second, an item-level ceiling: the same formal profiles explain only 3.3 to 3.6 percent of entropy variance and reach a median-split AUC of 0.606, too weak to identify high-disagreement items. Third, composition invariance: across the boundary, three high-powered preregistered contrasts on validated error shares and explanation-type shares (VariErr, LiTEx) all return null results. In this sample, formal semantic structure shifts how much annotators disagree by a small amount and does not detectably change what they disagree about. ChaosNLI-S/M consists of items selected for low original agreement, and every claim is conditioned on that scope. All analyses were preregistered in a version-controlled research log, whose audit trail, including one corrected interpretation rule, the paper discloses.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @chaosnli 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-much-human-label…] indexed:0 read:1min 2026-07-20 ·