{"slug": "how-much-human-label-variation-does-formal-semantic-structure-explain-group-and", "title": "How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI", "summary": "A new study measuring how much formal semantic structure explains human label variation in natural language inference finds that it accounts for only 3.3 to 3.6 percent of entropy variance, with a group-level effect showing non-purely-upward-monotone hypotheses have higher label entropy (Cliff's delta = -0.284). The analysis, conducted on 3,113 SNLI and MNLI items from ChaosNLI using a rule-based operator and monotonicity tagger, reveals that formal semantic structure shifts disagreement amounts by a small amount but does not detectably change what annotators disagree about.", "body_md": "arXiv:2607.15870v1 Announce Type: new\nAbstract: Human label variation in natural language inference is increasingly treated as signal rather than noise, but how much of it formal semantic structure explains has not been measured directly. We measure it on the 3,113 SNLI and MNLI items of ChaosNLI, using a rule-based operator and monotonicity tagger validated against MED (0.883 agreement at the edit site, 0.807 on the sentence-level summary our analyses consume), three preregistered analysis blocks, and full reporting of negative results. Three bounds emerge. First, a group-level boundary: hypotheses that are not purely upward monotone show reliably higher label entropy (Cliff's delta = -0.284), and rank-based tests defend the effect against operator-presence and length reductions, though a bounded-outcome sensitivity check weakens the regression form of the length defense. Second, an item-level ceiling: the same formal profiles explain only 3.3 to 3.6 percent of entropy variance and reach a median-split AUC of 0.606, too weak to identify high-disagreement items. Third, composition invariance: across the boundary, three high-powered preregistered contrasts on validated error shares and explanation-type shares (VariErr, LiTEx) all return null results. In this sample, formal semantic structure shifts how much annotators disagree by a small amount and does not detectably change what they disagree about. ChaosNLI-S/M consists of items selected for low original agreement, and every claim is conditioned on that scope. All analyses were preregistered in a version-controlled research log, whose audit trail, including one corrected interpretation rule, the paper discloses.", "url": "https://wpnews.pro/news/how-much-human-label-variation-does-formal-semantic-structure-explain-group-and", "canonical_source": "https://arxiv.org/abs/2607.15870", "published_at": "2026-07-20 04:00:00+00:00", "updated_at": "2026-07-20 13:41:53.457558+00:00", "lang": "en", "topics": ["natural-language-processing", "ai-research"], "entities": ["ChaosNLI", "SNLI", "MNLI", "MED", "VariErr", "LiTEx"], "alternates": {"html": "https://wpnews.pro/news/how-much-human-label-variation-does-formal-semantic-structure-explain-group-and", "markdown": "https://wpnews.pro/news/how-much-human-label-variation-does-formal-semantic-structure-explain-group-and.md", "text": "https://wpnews.pro/news/how-much-human-label-variation-does-formal-semantic-structure-explain-group-and.txt", "jsonld": "https://wpnews.pro/news/how-much-human-label-variation-does-formal-semantic-structure-explain-group-and.jsonld"}}