cd /news/large-language-models/can-a-mud-evaluate-llms-a-99-proof-o… · home topics large-language-models article
[ARTICLE · art-68798] src=cruciblebench.ai ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Can a MUD evaluate LLMs? A $99 proof of concept

CrucibleBench, a proof-of-concept evaluation framework that places large language models in a persistent MUD (multi-user dungeon) over 50 turns with hidden social objectives, found that a single LLM-judge component inside its scoring stack reordered the leaderboard by up to six positions while aggregate reliability statistics remained silent. The project tested 13 models across 650 runs at a total cost of $99.59, and the authors report every result under two scoring configurations, treating the divergence as the paper's most generalizable finding.

read5 min views1 publishedJul 22, 2026
Can a MUD evaluate LLMs? A $99 proof of concept
Image: source

CrucibleBench places language models in a persistent MUD, a text world where NPCs remember, trust accumulates, and mistakes leave traces, and scores what they do over 50 turns with hidden social objectives.

Status: Phase 1 proof-of-concept released · 13 models · 650 runs · $99.59 billed. Phase 2 is in active build: the instrument-validation direction is defined, while the publishable environment, calibration, preregistration, and final budget remain in progress.

Lateral thinking with withered technology #

Nintendo's Gunpei Yokoi used the phrase to describe a design philosophy: take mature, inexpensive, well-understood technology and use it in a new way. CrucibleBench applies it to AI evaluation.

Instead of photorealistic simulation or browser automation, we start with a MUD: a multi-user dungeon, the persistent text worlds of the early internet. Its constraints are the point: a limited command space makes hallucinated actions detectable, NPCs with trust and suspicion state give explicit social feedback, and within-run persistence means items taken stay taken and trust earned stays earned.

We did not choose a MUD because it is charming. We chose it because its constraints make behavior measurable.

Old constraints solve modern measurement problems #

Static benchmarks measure what models know in isolation. They do not measure how models behave where trust must be earned, information is gated by relationships, and blunt questioning raises suspicion. #

01

An enumerable action space

7 command types, 12 rooms, 14 items. Hallucinated actions and wrong-room interactions are detectable, and action efficiency is measurable.

02

Explicit social feedback

4 NPCs carry trust and suspicion state (0–100) that moves in response to dialogue: feedback a model can adapt to within a run, or fail to.

03

Within-run persistence

Items taken stay taken; trust earned stays earned. Every run leaves a complete, replayable transcript of exploration and planning.

The central finding is about measurement, not rankings #

A single LLM-judge component inside the scoring stack reordered the leaderboard by up to six positions, while every aggregate reliability statistic stayed silent. We report every result under two scoring configurations and treat the divergence as the paper's most generalizable finding.

Judge ablation reorders the top of the board

Two of four scored dimensions route through a dialogue classifier whose per-model agreement with an independent judge spans 21.7% to 84.8%, instability the aggregate κ = 0.04 never reveals. Removing the classifier-dependent dimensions shifts six rankings beyond scenario-sampling

noise (90% paired block bootstrap). The largest mover shares a model family with the classifier. Benchmarks that use LLM judges should report per-subject agreement and ranking stability under judge ablation, not aggregate reliability alone.

Model Full CM Δ
Claude Sonnet 4.6 #4 #1 ▲ 3
DeepSeek R1 #7 #2 ▲ 5
Grok 4 #12 #10 ▲ 2
GPT-5.4 #1 #5 ▼ 4
Gemini 3.1 Pro #3 #9 ▼ 6
Mistral Large 3 #10 #12 ▼ 2
Model Classifier-min. Full score Success $ / run
Claude Sonnet 4.6 4.04 3.89 24% $0.125
DeepSeek R1 4.00 3.85 22% $0.119
Claude Opus 4.6 3.93 3.93 30% $0.205
GPT-5.2 3.91 3.88 38% $0.113
GPT-5.4 3.88 4.07 68% $0.060
Qwen 3.5 397B 3.81 3.81 30% $0.017
Claude Haiku 4.5 3.80 3.88 34% $0.039
GPT-5.3 Chat 3.73 3.72 40% $0.095
Gemini 3.1 Pro 3.71 3.91 48% $0.339
Grok 4 3.61 3.48 32% $0.834
DeepSeek V3.2 3.60 3.61 24% $0.008
Mistral Large 3 3.44 3.69 40% $0.017
OLMo 3.1 32B 2.01 1.93 4% $0.005

Mean scores on a 1–5 rubric scale, sorted by classifier-minimized subtotal. 50 runs per model: 5 seeds × 2 objectives × 5 repetitions, temperature 0.3, billing-verified via OpenRouter. Rankings are exploratory; confidence intervals overlap substantially among the top eight. Full protocol, CIs, and

statistics in the whitepaper.

Failures you can read in the transcript #

Three failure modes, each detected algorithmically from state-machine telemetry, with no judge involved. Dialogue looping is the dominant mode for every model tested, frontier included.

Dialogue looping 14–66% of frontier runs

Eight or more talk commands at a single NPC in one run. The agent repeats a failed conversational approach instead of adapting: the persistent-world cousin of a support agent repeating itself.

Wrong-room interaction severe in floor model

A talk command answered by "no one here." Reveals lost world-state tracking, analogous to calling an API that is not in scope. Grok 4 was the only frontier model with meaningful incidence (12%).

Exploration paralysis selective, floor-dominant

Two or fewer rooms across twenty-plus turns, or five consecutive look commands. Information gathering that never becomes goal-directed action.

What this is and is not #

This is

  • A proof-of-concept for persistent-world behavioral evaluation.
  • A compact MUD with hidden social objectives and rule-based mechanics.
  • A way to surface measurable, interpretable failure modes.
  • A full artifact release: 650 transcripts, source, scoring code, and the complete billing export.

This is not

  • A validated measure of general social intelligence.
  • A definitive leaderboard of frontier models.
  • Yet predictive of real-world agent deployment outcomes.
  • A claim that LLM judges are useless (rather, evidence they need per-subject audits).

Phase 2 is where this becomes a benchmark. Help us build it. #

CrucibleBench is an independent research effort. Phase 2 is being built for calibration; a provisional low/base/high budget is published now, and the final allocation will follow pilot data and preregistration. There are three ways in.

fund it · provisional $3,500 envelope build it · environment, objectives, calibration run it · post-calibration pilot cohort

Questions, or interested in a private evaluation? Write to contact@cruciblebench.ai

── more in #large-language-models 4 stories · sorted by recency
── more on @cruciblebench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/can-a-mud-evaluate-l…] indexed:0 read:5min 2026-07-22 ·