I really appreciate that the whole setup — scenes, locked labels, instruction blocks, scoring script — is public down to the file. That’s rarer than it should be.
I’ve got a small tool built around exactly this question — is a categorical predictor actually tracking ground truth, or is the agreement number just an artifact of a skewed base rate. It grades that relationship (chi-square + Cramér’s V, A-F) rather than a single agreement coefficient. It gew out of validating classifiers against physical ground-truth data in a completely different field, but the math doesn’t know or care what the categories mean.
Happy to run your five raters against your locked labels and post the per-rater grades back here, alongside whatever your second annotator turns up …it could be a useful cross-check either way.
No strings, just curious whether it converges with what you find.
Hope the annotator search goes well.