{"slug": "experts-rise-where-llms-disagree-using-cross-model-disagreement-to-target-expert", "title": "Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation", "summary": "A study on arXiv (2609.26926v1) found that having experts label cases of cross-LLM disagreement with rationales produced the highest LLM-labeling accuracy at 64.9% against expert labels, beating an expert-revised codebook at 57.8%. The best Question Answering setting also outperformed the expert-revised codebook at 60.5%, while Codebook Verifying, in which experts edit LLM-generated revisions, was the third feedback method tested. The authors report that targeting expert attention via LLM disagreement can shorten codebook revision from months to days without sacrificing labeling performance.", "body_md": "arXiv:2609.26926v1 Announce Type: new \nAbstract: Large-scale text annotation brings expert insight to millions of documents, often through a codebook that AI annotators follow. Developing a robust codebook, however, takes months. Large language models (LLMs) could speed this process by applying an early codebook to the data, surfacing cases with strong LLM disagreement, and eliciting expert feedback to address them. We examined three ways experts can provide feedback for LLM codebook revision: (i) editing LLM-generated revisions driven by cross-LLM disagreement (Codebook Verifying), (ii) answering questions about LLM disagreements (Question Answering), and (iii) labeling disagreement cases with rationales (Rationale Labeling). Experiments on thousands of tutoring-session transcripts show that Rationale Labeling yielded the highest LLM-labeling accuracy (64.9%) against expert labels, outperforming the expert-revised codebook (57.8%). The best Question Answering setting also outperformed it (60.5%). Our work shows that LLMs can be used to strategically target expert attention, shortening months of codebook revision to days without sacrificing labeling performance.", "url": "https://wpnews.pro/news/experts-rise-where-llms-disagree-using-cross-model-disagreement-to-target-expert", "canonical_source": "https://arxiv.org/abs/2609.26926", "published_at": "2026-09-24 04:00:00+00:00", "updated_at": "2026-09-24 04:32:35.555390+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "natural-language-processing", "machine-learning"], "entities": ["arXiv", "LLMs"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/experts-rise-where-llms-disagree-using-cross-model-disagreement-to-target-expert", "markdown": "https://wpnews.pro/news/experts-rise-where-llms-disagree-using-cross-model-disagreement-to-target-expert.md", "text": "https://wpnews.pro/news/experts-rise-where-llms-disagree-using-cross-model-disagreement-to-target-expert.txt", "jsonld": "https://wpnews.pro/news/experts-rise-where-llms-disagree-using-cross-model-disagreement-to-target-expert.jsonld"}}