15:39
2026-07-22
cruciblebench.ai
large-language-models
Can a MUD evaluate LLMs? A $99 proof of concept
CrucibleBench, a proof-of-concept evaluation framework that places large language models in a persistent MUD (multi-user dungeon) over 50 turns with hidden social objectives, found that a single LLM-jโฆ