{"slug": "bias-audits-detect-bias-but-disagree-on-ranking-evidence-from-ten-instruments", "title": "Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models", "summary": "A study of ten extrinsic bias-audit instruments run over ten frontier models found that eight of the ten tools detect occupational gender bias with confidence intervals clear of zero, but cross-tool rank agreement is indistinguishable from chance (Kendall's W=0.07, p=0.83), meaning no single audit supports ranking one model against another. The paper, posted as arXiv:2609.15995v1, reports that two widely cited direct-probe benchmarks are saturated because frontier models now answer neutrally, and that a positive control with six deliberately weaker models restored within-tool reliability without restoring cross-tool ranking, indicating the tools measure different constructs. Bias direction also split by audit format: forced-choice decision tools mostly over-corrected toward women and toward working-class candidates in 273 of 278 hiring decisions, while free generation and default coreference stayed stereotype-congruent; raw responses, code, and analysis are available at https://github.com/williamguey/bias-audit-agreement.", "body_md": "arXiv:2609.15995v1 Announce Type: new \nAbstract: Emerging AI regulation mandates bias audits of high-risk systems, and audit scores are beginning to be used to rank models. Both uses assume different audit tools measure the same thing well enough to compare. We test that assumption directly, running ten extrinsic audit instruments over a shared panel of ten frontier models through one pooled inference gateway, first on occupational gender bias, then on age and socioeconomic status. Detection succeeds while ranking fails. Eight of ten tools detect bias with confidence intervals clear of zero; two widely cited direct-probe benchmarks are saturated because frontier models now answer neutrally. But cross-tool rank agreement is indistinguishable from chance (Kendall's W=0.07, p=0.83). A positive control with six deliberately weaker models separates two explanations: within-tool reliability recovers once the panel spans real capability gaps, yet cross-tool ranking never recovers, which points to the tools measuring different constructs rather than one construct noisily. Even the direction of bias splits by audit format: forced-choice decision tools mostly over-correct (toward women, and toward working-class candidates in 273 of 278 hiring decisions), while free generation and default coreference stay stereotype-congruent. The pattern replicates on socioeconomic status; an apparent ranking agreement on age dissolves under the paper's own tool-inclusion rules. The practical message: a single audit can detect bias and estimate its direction within its own operationalization, but no single audit supports ranking one model against another. All raw responses, code, and the analysis that recomputes every reported number from source are available at https://github.com/williamguey/bias-audit-agreement.", "url": "https://wpnews.pro/news/bias-audits-detect-bias-but-disagree-on-ranking-evidence-from-ten-instruments", "canonical_source": "https://arxiv.org/abs/2609.15995", "published_at": "2026-09-16 04:00:00+00:00", "updated_at": "2026-09-16 04:05:48.164797+00:00", "lang": "en", "topics": ["ai-safety", "ai-ethics", "ai-research", "ai-policy"], "entities": ["arXiv", "Kendall's W", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/bias-audits-detect-bias-but-disagree-on-ranking-evidence-from-ten-instruments", "markdown": "https://wpnews.pro/news/bias-audits-detect-bias-but-disagree-on-ranking-evidence-from-ten-instruments.md", "text": "https://wpnews.pro/news/bias-audits-detect-bias-but-disagree-on-ranking-evidence-from-ten-instruments.txt", "jsonld": "https://wpnews.pro/news/bias-audits-detect-bias-but-disagree-on-ranking-evidence-from-ten-instruments.jsonld"}}