02:09
2026-07-23
runtimewire.com
artificial-intelligence
CrucibleBench finds an LLM judge shifted rankings by six spots
Removing two classifier-dependent scoring dimensions from CrucibleBench shifted language model leaderboard positions by as many as six spots, researchers Benjamin Davis and Philip Mims reported. The eโฆ