01:08
2026-08-02
lesswrong.com
artificial-intelligence
MUD as AI Evaluation and LLM-judge distortion in ways aggregate Īŗ misses
An independent research group's CrucibleBench experiment found that LLM-judge-based evaluation rankings are highly sensitive to classifier components, with aggregate Īŗ on probe detection at 0.04 and pā¦