Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation
A study on arXiv (2609.26926v1) found that having experts label cases of cross-LLM disagreement with rationales produced the highest LLM-labeling accuracy at 64.9% against expert labels, beating an ex…