cd /news/large-language-models/observational-equivalence-of-llm-and… · home topics large-language-models article
[ARTICLE · art-136629] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Observational Equivalence of LLM and Human Annotation

A new arXiv paper (arXiv:2609.22133v1) reports that ten large language models, three human experts, and 165 crowdsourced workers independently classified the same texts from 14 peer-reviewed political science studies using identical codebooks, and found LLM annotation quality observationally equivalent to human coding, with LLMs agreeing with expert coders at rates comparable to those among experts themselves. The authors attribute the equivalence to ambiguity in texts and coding rules, noting that when LLMs disagree with experts, experts are also more likely to disagree with one another, and that clarifying coding rules reduces disagreement among both experts and sufficiently capable LLMs. The paper proposes using disagreement across LLMs to identify difficult cases and refine codebooks, and develops ambiguity-aware bounds for downstream inference when a unique annotation cannot be defined for every text.

by read1 min views1 publishedSep 22, 2026

arXiv:2609.22133v1 Announce Type: new Abstract: In this paper, we show that LLM and human coding are observationally equivalent in terms of annotation quality: recent LLMs agree with expert coders at rates comparable to those observed among experts themselves. We demonstrate this through replications of text-classification tasks from 14 peer-reviewed political science studies, in which ten LLMs, three human experts, and 165 crowdsourced workers independently classify the same texts using identical codebooks. We find that this equivalence is driven by ambiguity in the texts and coding rules. When LLMs disagree with experts, experts are also more likely to disagree with one another, and clarifying coding rules reduces disagreement among both experts and sufficiently capable LLMs. Thus, there is little empirical basis for preferring human coding on the basis of annotation quality alone, while LLMs offer substantial advantages in speed and cost. We therefore argue that the central challenge of text annotation is no longer choosing between human and machine coders, but developing coding rules that minimize ambiguity and accounting for the ambiguity that remains. To this end, we propose using disagreement across LLMs to identify difficult cases and refine codebooks, and we develop ambiguity-aware bounds for downstream inference when a unique annotation cannot be defined for every text.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/observational-equiva…] indexed:0 read:1min 2026-09-22 ·