00:00
2026-06-23
machinelearning.apple.com
large-language-models
Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
A new study finds that panels of nine large language model judges effectively provide only about two independent votes' worth of information, because the models make correlated errors on the same itemβ¦