{"slug": "meta-introduces-gamut-benchmark-to-measure-factual-completeness-in-ai", "title": "Meta introduces GAMUT benchmark to measure factual completeness in AI", "summary": "Meta AI researchers released GAMUT (Grounded Assessment of Multimodal Factuality), a benchmark that measures whether AI answers include all necessary facts rather than just individual accuracy, on July 21, 2026. The benchmark, containing 1,813 questions across ten domains with expert-verified rubrics, found that the top-performing model, Gemini 3.1 Pro, scored only 58.7%. Meta has made the benchmark resources publicly available on Hugging Face.", "body_md": "# Meta introduces GAMUT benchmark to measure factual completeness in AI\n\nThe new evaluation framework tests whether AI models actually tell you everything you need to know, not just whether they get individual facts right\n\nMeta AI researchers have released GAMUT, a benchmark designed to measure something most AI evaluations quietly ignore: whether an AI’s answer is actually complete. Not just accurate, not just fluent, but whether it includes all the facts that matter.\n\nThe benchmark, short for Grounded Assessment of Multimodal Factuality, was published as an arXiv paper on July 21, 2026. It introduces a structured rubric system that converts “did the AI say everything it should have” into binary, machine-gradable checklists.\n\n## What GAMUT actually does\n\nGAMUT tackles the problem of omission with a two-level meta-rubric framework. The system organizes required content hierarchically, meaning it doesn’t just list facts an answer should include. It structures them by importance and category, then translates those requirements into yes-or-no questions that language models can reliably grade.\n\nThe benchmark contains 1,813 questions rooted in real wearable imagery across ten distinct domains. Each question comes paired with expert-verified rubrics. Human experts defined what a complete answer looks like before any AI was tested against it.\n\n## The results are humbling\n\nMeta evaluated 14 different AI models against the GAMUT benchmark. The top performer was Gemini 3.1 Pro, which scored 58.7%.\n\nThe benchmark demonstrated strong discriminative ability across the models tested, meaning it could meaningfully separate better performers from worse ones rather than clustering everything together.\n\nMeta has made the benchmark resources publicly available through Hugging Face.\n\n## Why this matters beyond the lab\n\nThe multimodal aspect of GAMUT adds another layer. The benchmark’s questions are grounded in wearable imagery, meaning models need to interpret visual inputs and produce comprehensive textual responses.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/meta-introduces-gamut-benchmark-to-measure-factual-completeness-in-ai", "canonical_source": "https://cryptobriefing.com/meta-gamut-benchmark-ai-factual-completeness/", "published_at": "2026-07-24 11:49:33+00:00", "updated_at": "2026-07-24 12:10:16.028450+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-tools"], "entities": ["Meta AI", "GAMUT", "Gemini 3.1 Pro", "Hugging Face"], "alternates": {"html": "https://wpnews.pro/news/meta-introduces-gamut-benchmark-to-measure-factual-completeness-in-ai", "markdown": "https://wpnews.pro/news/meta-introduces-gamut-benchmark-to-measure-factual-completeness-in-ai.md", "text": "https://wpnews.pro/news/meta-introduces-gamut-benchmark-to-measure-factual-completeness-in-ai.txt", "jsonld": "https://wpnews.pro/news/meta-introduces-gamut-benchmark-to-measure-factual-completeness-in-ai.jsonld"}}