Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration
A study of 36 judge-examinee pairs across four benchmarks and six models found that a model's task accuracy strongly predicts its judging accuracy (Pearson r ≥ 0.90 on most models) and inversely predi…