{"slug": "chain-of-models-cross-model-auditing-for-bias-robust-llm-judges", "title": "Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges", "summary": "A new arXiv preprint (2607.28636v1) introduces Chain-of-Models (CoM), an automated audit pipeline in which a second model inspects a first model's reasoning trace to reduce cognitive biases in LLM judges. Testing 9 models from 6 families across 4 biases and 4 factual datasets, the authors found that auditor identity matters: Kimi-K2.5 is the strongest standalone model on several biases yet a weak auditor for Qwen2.5-72B's biased traces, and the best auditor is bias-specific (GPT-4o for bandwagon, authority, distraction; GLM-5 for sycophancy). Their per-bias auditor selection rule achieved 0.884 accuracy on biased slices, outperforming the strongest single fixed auditor (0.824) and the no-audit baseline (0.805).", "body_md": "arXiv:2607.28636v1 Announce Type: new\nAbstract: LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-driven debiasing, which is brittle across bias types, or human evaluation, which does not scale. We study \\emph{Chain-of-Models} (CoM), an automated audit pipeline in which a second model inspects the first model's reasoning trace before producing the final judgment. The key design question is whether the auditor should be the same model, a same-family model, or a different-family model. Across 9 models from 6 families, 4 cognitive biases, and 4 factual datasets, we find that auditor identity matters in two ways. First, standalone bias resistance does not predict audit effectiveness: Kimi-K2.5 is the strongest standalone model on several biases, yet is a weak auditor for Qwen2.5-72B's biased traces. Second, the best auditor is bias-specific: GPT-4o is strongest on bandwagon, authority, and distraction, while GLM-5 is strongest on sycophancy. We operationalize these findings with a per-bias auditor selection rule that, given the bias type, scores candidates along functional diversity, per-bias standalone resistance, and calibrated audit effectiveness. Under a calibration/test split, the selector reaches the highest accuracy across the four biased slices ($0.884$ vs.\\ $0.824$ for the strongest single fixed auditor and $0.805$ for the no-audit baseline). We release data, configurations, and an LLM-agent skill at https://anonymous.4open.science/r/chain-of-models-B585 .", "url": "https://wpnews.pro/news/chain-of-models-cross-model-auditing-for-bias-robust-llm-judges", "canonical_source": "https://arxiv.org/abs/2607.28636", "published_at": "2026-08-03 04:00:00+00:00", "updated_at": "2026-08-03 04:03:42.600702+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-research"], "entities": ["arXiv", "Kimi-K2.5", "Qwen2.5-72B", "GPT-4o", "GLM-5", "Chain-of-Models"], "alternates": {"html": "https://wpnews.pro/news/chain-of-models-cross-model-auditing-for-bias-robust-llm-judges", "markdown": "https://wpnews.pro/news/chain-of-models-cross-model-auditing-for-bias-robust-llm-judges.md", "text": "https://wpnews.pro/news/chain-of-models-cross-model-auditing-for-bias-robust-llm-judges.txt", "jsonld": "https://wpnews.pro/news/chain-of-models-cross-model-auditing-for-bias-robust-llm-judges.jsonld"}}