# Looking for one independent annotator: our LLM raters agree with each other at chance level and we can't tell why

> Source: <https://discuss.huggingface.co/t/looking-for-one-independent-annotator-our-llm-raters-agree-with-each-other-at-chance-level-and-we-cant-tell-why/180645#post_2>
> Published: 2026-09-21 14:42:14+00:00

I really appreciate that the whole setup — scenes, locked labels, instruction blocks, scoring script — is public down to the file. That’s rarer than it should be.

I’ve got a small tool built around exactly this question — is a categorical predictor actually tracking ground truth, or is the agreement number just an artifact of a skewed base rate. It grades that relationship (chi-square + Cramér’s V, A-F) rather than a single agreement coefficient. It gew out of validating classifiers against physical ground-truth data in a completely different field, but the math doesn’t know or care what the categories mean.

Happy to run your five raters against your locked labels and post the per-rater grades back here, alongside whatever your second annotator turns up …it could be a useful cross-check either way.

No strings, just curious whether it converges with what you find.

Hope the annotator search goes well.
