cd/entity/CrucibleBenchยท homeโ€บ entitiesโ€บ CrucibleBench
grep -l @cruciblebench /news/*.json | wc -l โ†’ 1

CrucibleBench

mentions 1 type Organization feed RSS

// recent coverage 1 mentions

15:39
2026-07-22
cruciblebench.ai
large-language-models

Can a MUD evaluate LLMs? A $99 proof of concept

CrucibleBench, a proof-of-concept evaluation framework that places large language models in a persistent MUD (multi-user dungeon) over 50 turns with hidden social objectives, found that a single LLM-jโ€ฆ

// co-occurs with top 7 entities
// topics top 2 topics