01:08
2026-08-02
lesswrong.com
artificial-intelligence
MUD as AI Evaluation and LLM-judge distortion in ways aggregate κ misses
An independent research group's CrucibleBench experiment found that LLM-judge-based evaluation rankings are highly sensitive to classifier components, with aggregate κ on probe detection at 0.04 and p…