{"slug": "reliable-but-design-sensitive-instrument-uncertainty-in-llm-annotation", "title": "Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation", "summary": "A study of seven LLMs, 12 task designs, three independent runs and 3,000 tweets labeled for offensive language and hate speech found that repeating the same model and task design produced high agreement (median Fleiss' κ = 0.91), but agreement fell when the task design changed for the same tweets (median Cohen's κ = 0.76). Task design and model choice increased the variance of estimated prevalence by factors of 76.7 for offensive language and 110.6 for hate speech compared with sampling variance alone, with variation across LLM task designs reaching 560-572 basis points versus 270-331 basis points across five human instrument versions. The authors term this \"instrument uncertainty\" and report that confidence scores did not solve it, tracking repeated model outputs more closely than agreement with human labels, while grouping six tweets in one prompt lowered mean offensive-language confidence by 660 basis points.", "body_md": "arXiv:2609.35824v1 Announce Type: new \nAbstract: Large language models (LLMs) can give reliable labels under one setup yet change those labels when researchers make other reasonable design choices. We tested seven LLMs, 12 task designs, three independent runs, and 3,000 tweets labeled for offensive language and hate speech. Repeating the same model and task design produced high agreement (median Fleiss' $\\kappa = 0.91$). Agreement fell when we changed the task design for the same tweets (median Cohen's $\\kappa = 0.76$). Task design and model choice increased the variance of estimated prevalence by factors of 76.7 for offensive language and 110.6 for hate speech compared with sampling variance alone. Variation across LLM task designs reached 560-572 basis points, compared with 270-331 basis points across five human instrument versions. Confidence scores did not solve this problem. They tracked repeated model outputs more closely than agreement with human labels, and grouping six tweets in one prompt lowered mean offensive-language confidence by 660 basis points. We call the variation caused by task design and model choice instrument uncertainty. Researchers can measure it only by comparing reasonable task designs. Repeating one setup or relying on confidence scores cannot replace that test.", "url": "https://wpnews.pro/news/reliable-but-design-sensitive-instrument-uncertainty-in-llm-annotation", "canonical_source": "https://arxiv.org/abs/2609.35824", "published_at": "2026-09-30 04:00:00+00:00", "updated_at": "2026-09-30 04:20:39.027338+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "natural-language-processing", "ai-safety"], "entities": ["arXiv", "Fleiss' kappa", "Cohen's kappa"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/reliable-but-design-sensitive-instrument-uncertainty-in-llm-annotation", "markdown": "https://wpnews.pro/news/reliable-but-design-sensitive-instrument-uncertainty-in-llm-annotation.md", "text": "https://wpnews.pro/news/reliable-but-design-sensitive-instrument-uncertainty-in-llm-annotation.txt", "jsonld": "https://wpnews.pro/news/reliable-but-design-sensitive-instrument-uncertainty-in-llm-annotation.jsonld"}}