{"slug": "uncertaintygym-a-benchmark-for-llm-epistemic-calibration-and-uncertainty", "title": "UncertaintyGym: A Benchmark for LLM Epistemic Calibration and Uncertainty Expression", "summary": "Muse Ltd. released UncertaintyGym, a benchmark for evaluating large language models' epistemic calibration and uncertainty expression, available on Hugging Face under the Apache-2.0 license. The benchmark measures whether models explicitly express uncertainty, request missing context, or reject false premises across four categories: solvable, under-specified, false premise, and inherently unknowable. Initial baseline results show LiquidAI's LFM2.5-2.6B model achieved a Meta Cognitive Calibration Score of 45.0%, with a hallucination rate of 73.3%.", "body_md": "Hi Hugging Face team,\n\nWe would like to submit UncertaintyGym for official benchmark indexing and Evaluation Hub integration.\n\nOverview:\n\nUncertaintyGym evaluates an LLM’s meta-cognitive calibration—measuring whether models explicitly express uncertainty (“I don’t know”), request missing context, or reject false premises instead of hallucinating.\n\nTaxonomy:\n\n- Category A (Solvable): Standard factual queries (control set).\n\n- Category B (Under-specified): Queries lacking necessary context that require disambiguation.\n\n- Category C (False Premise): Queries with impossible premises that require rejection.\n\n- Category D (Inherently Unknowable): Unsolved conjectures, lost history, and future events requiring explicit declaration of unknowability.\n\nTechnical Specifications:\n\n- Repository: [Muse-Ltd/UncertaintyGym · Datasets at Hugging Face](https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym)\n\n- Configuration: eval.yaml and lighteval_task.py included for direct LightEval / LM-Evaluation-Harness integration.\n\n- Primary Metric: Meta Cognitive Calibration Score (MCS) and Unanswerable Hallucination Rate.\n\n- License: Apache-2.0\n\nInitial Baseline:\n\n- LiquidAI / LFM2.5-2.6B: 45.0% MCS (Cat A: 100%, Cat B: 20%, Cat C: 0%, Cat D: 60%, Hallucination Rate: 73.3%)\n\nPlease let us know if any additional metadata is required for the benchmark badge.\n\nThank you.", "url": "https://wpnews.pro/news/uncertaintygym-a-benchmark-for-llm-epistemic-calibration-and-uncertainty", "canonical_source": "https://discuss.huggingface.co/t/uncertaintygym-a-benchmark-for-llm-epistemic-calibration-and-uncertainty-expression/178647#post_1", "published_at": "2026-08-13 19:26:31+00:00", "updated_at": "2026-08-13 19:45:02.017215+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-ethics", "ai-tools"], "entities": ["Muse Ltd.", "UncertaintyGym", "Hugging Face", "LiquidAI", "LFM2.5-2.6B"], "alternates": {"html": "https://wpnews.pro/news/uncertaintygym-a-benchmark-for-llm-epistemic-calibration-and-uncertainty", "markdown": "https://wpnews.pro/news/uncertaintygym-a-benchmark-for-llm-epistemic-calibration-and-uncertainty.md", "text": "https://wpnews.pro/news/uncertaintygym-a-benchmark-for-llm-epistemic-calibration-and-uncertainty.txt", "jsonld": "https://wpnews.pro/news/uncertaintygym-a-benchmark-for-llm-epistemic-calibration-and-uncertainty.jsonld"}}