cd /news/large-language-models/reliable-but-design-sensitive-instru… · home › topics › large-language-models › article
[ARTICLE · art-142252] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation

A study of seven LLMs, 12 task designs, three independent runs and 3,000 tweets labeled for offensive language and hate speech found that repeating the same model and task design produced high agreement (median Fleiss' κ = 0.91), but agreement fell when the task design changed for the same tweets (median Cohen's κ = 0.76). Task design and model choice increased the variance of estimated prevalence by factors of 76.7 for offensive language and 110.6 for hate speech compared with sampling variance alone, with variation across LLM task designs reaching 560-572 basis points versus 270-331 basis points across five human instrument versions. The authors term this "instrument uncertainty" and report that confidence scores did not solve it, tracking repeated model outputs more closely than agreement with human labels, while grouping six tweets in one prompt lowered mean offensive-language confidence by 660 basis points.

by read1 min views1 publishedSep 30, 2026

arXiv:2609.35824v1 Announce Type: new Abstract: Large language models (LLMs) can give reliable labels under one setup yet change those labels when researchers make other reasonable design choices. We tested seven LLMs, 12 task designs, three independent runs, and 3,000 tweets labeled for offensive language and hate speech. Repeating the same model and task design produced high agreement (median Fleiss' $\kappa = 0.91$). Agreement fell when we changed the task design for the same tweets (median Cohen's $\kappa = 0.76$). Task design and model choice increased the variance of estimated prevalence by factors of 76.7 for offensive language and 110.6 for hate speech compared with sampling variance alone. Variation across LLM task designs reached 560-572 basis points, compared with 270-331 basis points across five human instrument versions. Confidence scores did not solve this problem. They tracked repeated model outputs more closely than agreement with human labels, and grouping six tweets in one prompt lowered mean offensive-language confidence by 660 basis points. We call the variation caused by task design and model choice instrument uncertainty. Researchers can measure it only by comparing reasonable task designs. Repeating one setup or relying on confidence scores cannot replace that test.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reliable-but-design-…] indexed:0 read:1min 2026-09-30 · —