{"slug": "looking-for-one-independent-annotator-our-llm-raters-agree-with-each-other-at-we", "title": "Looking for one independent annotator: our LLM raters agree with each other at chance level and we can't tell why", "summary": "A commenter offered to independently grade five LLM raters against a project's locked labels using chi-square and Cramér's V scoring, after the project reported that its raters agree with each other at chance level. The commenter said the tool, developed for validating classifiers against physical ground-truth data, would produce per-rater grades from A to F to cross-check the project's own second annotator. The project's scenes, locked labels, instruction blocks, and scoring script are public.", "body_md": "I really appreciate that the whole setup — scenes, locked labels, instruction blocks, scoring script — is public down to the file. That’s rarer than it should be.\n\nI’ve got a small tool built around exactly this question — is a categorical predictor actually tracking ground truth, or is the agreement number just an artifact of a skewed base rate. It grades that relationship (chi-square + Cramér’s V, A-F) rather than a single agreement coefficient. It gew out of validating classifiers against physical ground-truth data in a completely different field, but the math doesn’t know or care what the categories mean.\n\nHappy to run your five raters against your locked labels and post the per-rater grades back here, alongside whatever your second annotator turns up …it could be a useful cross-check either way.\n\nNo strings, just curious whether it converges with what you find.\n\nHope the annotator search goes well.", "url": "https://wpnews.pro/news/looking-for-one-independent-annotator-our-llm-raters-agree-with-each-other-at-we", "canonical_source": "https://discuss.huggingface.co/t/looking-for-one-independent-annotator-our-llm-raters-agree-with-each-other-at-chance-level-and-we-cant-tell-why/180645#post_2", "published_at": "2026-09-21 14:42:14+00:00", "updated_at": "2026-09-21 14:52:50.565763+00:00", "lang": "en", "topics": ["large-language-models", "ai-research"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/looking-for-one-independent-annotator-our-llm-raters-agree-with-each-other-at-we", "markdown": "https://wpnews.pro/news/looking-for-one-independent-annotator-our-llm-raters-agree-with-each-other-at-we.md", "text": "https://wpnews.pro/news/looking-for-one-independent-annotator-our-llm-raters-agree-with-each-other-at-we.txt", "jsonld": "https://wpnews.pro/news/looking-for-one-independent-annotator-our-llm-raters-agree-with-each-other-at-we.jsonld"}}