arXiv:2610.07115v1 Announce Type: new Abstract: The order in which candidate responses are presented can change an LLM judge's verdict. Detecting such a position flip ordinarily requires judging each pair in both orders, which doubles the number of judgments. We investigate whether residual stream activations recorded immediately before the initial verdict can predict a flip. We use nested grouped cross-validation to evaluate regularized linear probes on 534 JudgeBench pairs for three Qwen3 judges and Llama-3.1-8B. The linear probes achieve AUROCs of .621-.850 and outperform a combined baseline that uses verbalized confidence, verdict-label logits, response lengths, and the judge's initial choice by .062-.113 AUROC. Linear probes trained on JudgeBench and then frozen achieve AUROCs of .685-.853 on 1,802 MT-Bench comparisons without MT-Bench fitting or recalibration. These results show that pre-verdict activations support prediction of susceptibility to candidate order and outperform the non-activation predictors evaluated here.
Will the Judge Flip? Predicting Position-Sensitive LLM Judgments from Residual Stream Activations
Linear probes on residual stream activations recorded immediately before an LLM judge's initial verdict can predict whether that verdict will flip when candidate responses are reordered, according to an arXiv paper (2610.07115v1). Evaluated on 534 JudgeBench pairs across three Qwen3 judges and Llama-3.1-8B using nested grouped cross-validation, the probes reached AUROCs of .621-.850, beating a combined baseline of verbalized confidence, verdict-label logits, response lengths, and the judge's initial choice by .062-.113 AUROC. Probes trained on JudgeBench and then frozen scored .685-.853 AUROC on 1,802 MT-Bench comparisons without MT-Bench fitting or recalibration.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.