cd /news/large-language-models/evaluating-llms-via-muds-a-99-experi… · home topics large-language-models article
[ARTICLE · art-71038] src=promptcube3.com ↗ pub= topic=large-language-models verified=true sentiment=↓ negative

Evaluating LLMs via MUDs: A $99 Experiment

A $99 experiment evaluating LLMs via MUDs found that removing two dimensions reliant on an LLM classifier caused one frontier model to drop six spots on the leaderboard, with agreement rates between judges swinging from 85% to 22% and a probe detection kappa of 0.04, highlighting systemic noise in judge-based evaluation. The study, which ran 50 runs per model with no human raters, underscores that a judge is only as good as its consistency.

read1 min views1 publishedJul 23, 2026
Evaluating LLMs via MUDs: A $99 Experiment
Image: Promptcube3 (auto-discovered)

We measured models across four behavioral dimensions. The results were jarring: when we stripped away the two dimensions that relied on an LLM classifier, one frontier model plummeted six spots on the leaderboard. Even more concerning was the agreement rate between two different judges, which swung wildly from 85% down to 22% depending on the model. The aggregate kappa for probe detection was a dismal 0.04, meaning the instrument was incredibly noisy. Interestingly, the model most affected by this noise belonged to the same family as the classifier.

This highlights a critical flaw in many current AI workflows and benchmarks: the "judge" model often introduces systemic noise or bias that skews the perceived performance of the "student" model.

This wasn't a full-scale validation, but rather a deep dive into the possibilities and pitfalls of judge-based evaluation. The limitations were plenty:

Sample size: Only 50 runs per model.Confidence: Overlapping CIs among top performers.Validation: Zero human raters involved.Scale: A very small game environment.

For anyone building their own LLM agent or benchmarking a custom prompt engineering setup, this is a reminder that your judge is only as good as its consistency.

The full dataset, paper, and billing exports are available here:

https://doi.org/10.5281/zenodo.21386663

Next Hale: A New Concurrent Systems Language →

── more in #large-language-models 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evaluating-llms-via-…] indexed:0 read:1min 2026-07-23 ·