04:00
2026-09-14
arxiv.org
large-language-models
Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration
A study of 36 judge-examinee pairs across four benchmarks and six models found that a model's task accuracy strongly predicts its judging accuracy (Pearson r β₯ 0.90 on most models) and inversely prediβ¦