We measured models across four behavioral dimensions. The results were jarring: when we stripped away the two dimensions that relied on an LLM classifier, one frontier model plummeted six spots on the leaderboard. Even more concerning was the agreement rate between two different judges, which swung wildly from 85% down to 22% depending on the model. The aggregate kappa for probe detection was a dismal 0.04, meaning the instrument was incredibly noisy. Interestingly, the model most affected by this noise belonged to the same family as the classifier.
This highlights a critical flaw in many current AI workflows and benchmarks: the "judge" model often introduces systemic noise or bias that skews the perceived performance of the "student" model.
This wasn't a full-scale validation, but rather a deep dive into the possibilities and pitfalls of judge-based evaluation. The limitations were plenty:
Sample size: Only 50 runs per model.Confidence: Overlapping CIs among top performers.Validation: Zero human raters involved.Scale: A very small game environment.
For anyone building their own LLM agent or benchmarking a custom prompt engineering setup, this is a reminder that your judge is only as good as its consistency.
The full dataset, paper, and billing exports are available here:
https://doi.org/10.5281/zenodo.21386663
Next Hale: A New Concurrent Systems Language →