Looking for one independent annotator: our LLM raters agree with each other at chance level and we can't tell why A commenter offered to independently grade five LLM raters against a project's locked labels using chi-square and Cramér's V scoring, after the project reported that its raters agree with each other at chance level. The commenter said the tool, developed for validating classifiers against physical ground-truth data, would produce per-rater grades from A to F to cross-check the project's own second annotator. The project's scenes, locked labels, instruction blocks, and scoring script are public. I really appreciate that the whole setup — scenes, locked labels, instruction blocks, scoring script — is public down to the file. That’s rarer than it should be. I’ve got a small tool built around exactly this question — is a categorical predictor actually tracking ground truth, or is the agreement number just an artifact of a skewed base rate. It grades that relationship chi-square + Cramér’s V, A-F rather than a single agreement coefficient. It gew out of validating classifiers against physical ground-truth data in a completely different field, but the math doesn’t know or care what the categories mean. Happy to run your five raters against your locked labels and post the per-rater grades back here, alongside whatever your second annotator turns up …it could be a useful cross-check either way. No strings, just curious whether it converges with what you find. Hope the annotator search goes well.