cd /news/large-language-models/experts-rise-where-llms-disagree-usi… · home topics large-language-models article
[ARTICLE · art-138827] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation

A study on arXiv (2609.26926v1) found that having experts label cases of cross-LLM disagreement with rationales produced the highest LLM-labeling accuracy at 64.9% against expert labels, beating an expert-revised codebook at 57.8%. The best Question Answering setting also outperformed the expert-revised codebook at 60.5%, while Codebook Verifying, in which experts edit LLM-generated revisions, was the third feedback method tested. The authors report that targeting expert attention via LLM disagreement can shorten codebook revision from months to days without sacrificing labeling performance.

by read1 min views1 publishedSep 24, 2026

arXiv:2609.26926v1 Announce Type: new Abstract: Large-scale text annotation brings expert insight to millions of documents, often through a codebook that AI annotators follow. Developing a robust codebook, however, takes months. Large language models (LLMs) could speed this process by applying an early codebook to the data, surfacing cases with strong LLM disagreement, and eliciting expert feedback to address them. We examined three ways experts can provide feedback for LLM codebook revision: (i) editing LLM-generated revisions driven by cross-LLM disagreement (Codebook Verifying), (ii) answering questions about LLM disagreements (Question Answering), and (iii) labeling disagreement cases with rationales (Rationale Labeling). Experiments on thousands of tutoring-session transcripts show that Rationale Labeling yielded the highest LLM-labeling accuracy (64.9%) against expert labels, outperforming the expert-revised codebook (57.8%). The best Question Answering setting also outperformed it (60.5%). Our work shows that LLMs can be used to strategically target expert attention, shortening months of codebook revision to days without sacrificing labeling performance.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/experts-rise-where-l…] indexed:0 read:1min 2026-09-24 ·