MUD as AI Evaluation and LLM-judge distortion in ways aggregate κ misses
An independent research group's CrucibleBench experiment found that LLM-judge-based evaluation rankings are highly sensitive to classifier components, with aggregate κ on probe detection at 0.04 and per-model agreement r…