Observational Equivalence of LLM and Human Annotation A new arXiv paper (arXiv:2609.22133v1) reports that ten large language models, three human experts, and 165 crowdsourced workers independently classified the same texts from 14 peer-reviewed political science studies using identical codebooks, and found LLM annotation quality observationally equivalent to human coding, with LLMs agreeing with expert coders at rates comparable to those among experts themselves. The authors attribute the equivalence to ambiguity in texts and coding rules, noting that when LLMs disagree with experts, experts are also more likely to disagree with one another, and that clarifying coding rules reduces disagreement among both experts and sufficiently capable LLMs. The paper proposes using disagreement across LLMs to identify difficult cases and refine codebooks, and develops ambiguity-aware bounds for downstream inference when a unique annotation cannot be defined for every text. arXiv:2609.22133v1 Announce Type: new Abstract: In this paper, we show that LLM and human coding are observationally equivalent in terms of annotation quality: recent LLMs agree with expert coders at rates comparable to those observed among experts themselves. We demonstrate this through replications of text-classification tasks from 14 peer-reviewed political science studies, in which ten LLMs, three human experts, and 165 crowdsourced workers independently classify the same texts using identical codebooks. We find that this equivalence is driven by ambiguity in the texts and coding rules. When LLMs disagree with experts, experts are also more likely to disagree with one another, and clarifying coding rules reduces disagreement among both experts and sufficiently capable LLMs. Thus, there is little empirical basis for preferring human coding on the basis of annotation quality alone, while LLMs offer substantial advantages in speed and cost. We therefore argue that the central challenge of text annotation is no longer choosing between human and machine coders, but developing coding rules that minimize ambiguity and accounting for the ambiguity that remains. To this end, we propose using disagreement across LLMs to identify difficult cases and refine codebooks, and we develop ambiguity-aware bounds for downstream inference when a unique annotation cannot be defined for every text.