{"slug": "epoch-ai-audits-15-benchmarks-finds-nine-flawed", "title": "Epoch AI audits 15 benchmarks, finds nine flawed", "summary": "Epoch AI launched Benchmark Reviews on September 17th and classified nine of its first 15 reviewed benchmark versions as Flawed, four as Verified, and two as Not Enough Info, flagging defects in tests including SWE-bench Verified and Humanity's Last Exam. The research institute led by co-founder Jaime Sevilla assigns a Flawed verdict when at least 20% of an inspected sample contains errors or one issue corrupts grading at scale, with reviews completed between August 10th and September 12th. Epoch AI says the audits show some headline leaderboard scores reflect broken graders and ambiguous tasks rather than model capability, and the program adds a public quality-control layer to its registry of 85 benchmarks.", "body_md": "# Epoch AI audits 15 benchmarks, finds nine flawed\n\n**The first batch flags defects in SWE-bench Verified, Humanity's Last Exam and other tests used to compare frontier models.**\n\n        By [RuntimeWire Staff](https://runtimewire.com/author/runtimewire-staff)\n        · Published \n\nPrimary source: [Epoch AI](https://x.com/EpochAIResearch/status/2100704765332394255)\n\n## Why it matters\n\nAI leaderboards guide product claims, deployment decisions and policy. Epoch AI's audits show that some headline scores reflect broken graders and ambiguous tasks rather than model capability.\n\n[Epoch AI](https://epoch.ai/?ref=runtimewire), the research institute co-founded and led by [Jaime Sevilla (@Jsevillamol)](https://x.com/Jsevillamol?ref=runtimewire), launched Benchmark Reviews on September 17th and classified nine of its first 15 reviewed benchmark versions as Flawed. Four received Verified labels, while Epoch AI lacked enough information to judge the remaining two.\n\nThe result gives Sevilla a new way to pursue the thesis behind Epoch AI: AI progress requires better measurement than the industry's usual mixture of leaderboards, selective model cards and launch-day claims. He started Epoch AI as a volunteer data-collection project in 2021, then built it into a research institute after a [2022 study of machine-learning compute trends](https://arxiv.org/abs/2202.05924?ref=runtimewire) established the group's public profile.\n\nSevilla studied mathematics and computer engineering at the Complutense University of Madrid, worked as a deep-learning engineer and began doctoral research on explainable AI before focusing on forecasting technological progress. Epoch AI says the founding group was frustrated that a consequential industry was still being measured through \"hype and vibes.\" Benchmark Reviews applies that complaint to the tests now used as evidence that models can code, call tools, answer expert questions or complete agentic workflows.\n\nEpoch AI [announced the initiative in a thread on X](https://x.com/EpochAIResearch/status/2100704765332394255?ref=runtimewire). The initial reviews were completed between August 10th and September 12th, making Thursday's launch the publication of an accumulated audit program rather than 15 reviews performed at once.\n\n### A benchmark for benchmarks\n\nEpoch AI assigns each reviewed benchmark version one of three verdicts. Verified means the results can broadly be interpreted as the benchmark creator describes them. Flawed means Epoch AI found a substantive score-affecting defect. Not Enough Info means the tasks, scoring logic or evaluation settings were too inaccessible to support either conclusion.\n\nUnder Epoch AI's [published methodology](https://epoch.ai/data/benchmark-reviews-documentation/methodology?ref=runtimewire), a benchmark normally receives a Flawed verdict when at least 20% of an inspected sample contains errors, or when one issue corrupts grading at scale. The review also checks for changing tests without version updates, restrictive evaluation setups and model-specific advantages caused by unequal scaffolds or compute budgets.\n\nFor benchmarks with more than 50 available tasks, Epoch AI generally samples 50, expanding to 100 when the observed error rate falls between 15% and 25%. It examines every available task when a benchmark has 50 or fewer. Those sampling rules make the verdicts reproducible, while leaving room for defects outside the inspected set. Epoch AI stops a review once it has gathered enough evidence for a Flawed label, so those writeups should be read as a minimum case rather than a complete defect inventory.\n\nGreg Burnham, Epoch AI's head of benchmarks, oversees a function that has moved from running evaluations toward inspecting the machinery underneath them. Burnham studied mathematics at Princeton and previously worked at Elemental Cognition and Bridgewater Associates. The new program turns that work into a public quality-control layer across Epoch AI's [registry of 85 benchmarks](https://epoch.ai/benchmarks/search?reviewed=verified&reviewed=flawed&reviewed=not-enough-info&ref=runtimewire).\n\n### What failed\n\nSome of the findings cut directly into benchmarks that model developers use to substantiate product claims.\n\nEpoch AI's [review of Humanity's Last Exam](https://epoch.ai/benchmarks/hle/review?ref=runtimewire) found substantial accuracy-altering errors in 22 of 48 sampled questions, or 46%. Twelve questions were judged impossible to answer correctly as written. Other problems could reject valid answers or accept incorrect ones. One question requested a four-point discrete Fourier transform while supplying eight entries; another answer key pointed to the wrong multiple-choice option.\n\nIn the [Berkeley Function Calling Leaderboard v4 review](https://epoch.ai/benchmarks/berkeley-function-calling-leaderboard/review?ref=runtimewire), Epoch AI found potential accuracy defects in 24 of 50 sampled tasks. The benchmark tests whether models select tools and supply the correct arguments. Epoch AI found cases where correct behavior could be penalized, including a task that expected a model to add a stock already present on a watchlist and a web-search question whose answer had become stale.\n\n[DeepSWE v1.1](https://epoch.ai/benchmarks/deepswe/review?ref=runtimewire) crossed the threshold narrowly but concretely: Epoch AI confirmed false negatives in at least 23 of 113 tasks, or 20.3%. In many cases, the grader discarded changes that an agent made to test files, then failed because surviving code still depended on those changes. The model could produce a reasonable patch and lose points because the verification system broke the submission.\n\nThe driest finding belongs to [SWE-bench Verified](https://epoch.ai/benchmarks/swe-bench-verified/review?ref=runtimewire), whose name did not save it from a Flawed verdict. Epoch AI relied partly on an OpenAI audit that examined 27.6% of its 500 tasks and found flawed tests in 59.4% of the audited subset. Epoch AI also pointed to contamination risk because the benchmark uses public repositories likely to appear in model training data.\n\nVerified does not mean spotless. Epoch AI found score-affecting defects in five of 50 sampled questions from [SimpleQA Verified](https://epoch.ai/benchmarks/simple-qa-verified/review?ref=runtimewire), below its 20% cutoff. The review also warns that the score partly measures a model's willingness to guess because abstentions count as wrong. A leaderboard can therefore reward confidence alongside factual recall.\n\n### The denominator needs its own warning label\n\nNine Flawed verdicts out of 15 reviews is a sharp launch statistic. It is not evidence that 60% of AI benchmarks are defective.\n\nEpoch AI says the [initial set](https://epoch.ai/data/benchmark-reviews-documentation/included-benchmarks?ref=runtimewire) was assembled to cover different domains and error types. That is a curated sample rather than a random draw from the broader registry. The launch tally establishes that serious defects exist in several prominent benchmarks. It cannot establish their prevalence across the entire evaluation market.\n\nThe verdicts also apply to named versions. Benchmark creators can publish corrected versions and seek another review. Epoch AI says it will preserve the original verdict as a record of the reviewed snapshot, giving developers an incentive to version changes rather than silently altering tasks or scoring rules.\n\nThat version discipline matters because benchmark scores increasingly shape product marketing, technical roadmaps and judgments about whether models are ready for deployment. A few percentage points can reorder a leaderboard. Epoch AI's reviews show that those points may come from stale answers, hidden grader behavior, uneven scaffolds or ambiguous questions rather than a meaningful capability difference.\n\n### Epoch AI writes a conflict rule after FrontierMath\n\nEpoch AI will not review benchmarks it created, citing conflicts of interest, and says it may invite external audits instead. The rule carries weight because Sevilla has already had to tighten Epoch AI's standards around disclosure.\n\nIn January 2025, Epoch AI [acknowledged](https://epoch.ai/latest/openai-and-frontiermath?ref=runtimewire) that OpenAI had commissioned the 300-question core of FrontierMath and could access its questions and solutions, apart from a 50-question holdout set. Epoch AI said its initial communication about the relationship was insufficient. Sevilla later wrote that future benchmark projects would retain ownership, provide more equitable access and disclose funding relationships proactively.\n\nThe no-self-review policy does not claim that Epoch AI's benchmarks are clean. Its documentation points to an audit that found errors in 42% of FrontierMath v1 problems. Excluding its own work preserves a clearer boundary: Epoch AI can build benchmarks or judge them, but it will not give its own tests a Verified badge.\n\nBenchmark Reviews now has to prove that boundary works in practice. Epoch AI says it will prioritize influential and safety-relevant benchmarks, notify creators before publication, publish their responses if requested and consider updated versions for fresh reviews. The program's value will come from maintaining that process when a Flawed label lands on a benchmark backed by a large lab, a commercial evaluator or a well-connected research group.\n\nFor Sevilla, the product is an institutional bet that benchmark quality can become a monitored public record rather than an occasional postmortem. The first batch already shows why the audit layer is needed: AI models are being ranked with instruments that sometimes grade correct work as failure, accept bad answers or test a different capability from the one printed on the label.", "url": "https://wpnews.pro/news/epoch-ai-audits-15-benchmarks-finds-nine-flawed", "canonical_source": "https://runtimewire.com/article/epoch-ai-benchmark-reviews-nine-flawed", "published_at": "2026-09-18 06:57:43+00:00", "updated_at": "2026-09-18 07:25:48.931526+00:00", "lang": "en", "topics": ["ai-research", "ai-safety", "large-language-models", "ai-agents", "ai-policy"], "entities": ["Epoch AI", "Jaime Sevilla", "Greg Burnham", "SWE-bench Verified", "Humanity's Last Exam", "Benchmark Reviews", "Complutense University of Madrid", "Princeton"], "alternates": {"html": "https://wpnews.pro/news/epoch-ai-audits-15-benchmarks-finds-nine-flawed", "markdown": "https://wpnews.pro/news/epoch-ai-audits-15-benchmarks-finds-nine-flawed.md", "text": "https://wpnews.pro/news/epoch-ai-audits-15-benchmarks-finds-nine-flawed.txt", "jsonld": "https://wpnews.pro/news/epoch-ai-audits-15-benchmarks-finds-nine-flawed.jsonld"}}