Can a Buyer Reproduce a Vendor Benchmark Row? An audit of 42 benchmark rows published by nine vendors — DeepSeek, Google, xAI, Z.ai, Anthropic, OpenAI, Moonshot AI, NVIDIA, and Mistral AI — on pages dated on or before August 14, 2026, found that zero rows were both fully disclosed and independently confirmed. Of the 41 rows scored, only 10 name an evaluation harness, 16 state a reasoning-effort setting, and 29 are missing at least one applicable disclosure field. The audit, which checked disclosure on all rows but attempted independent locatability on only 15, highlights that xAI's own page reports Grok 4.6 at 26% on Terminal-Bench v3.0 while an independent evaluator measured 88.4% on v2.1, and that DeepSWE v1.1 scores from four vendors fell within the independent leaderboard's uncertainty band. Vendor benchmark reproducibility has a simple test: could a buyer, starting from the vendor’s own launch page, work out exactly which test produced the number in the table? We collected 42 benchmark rows that nine vendors — DeepSeek, Google, xAI, Z.ai, Anthropic, OpenAI, Moonshot AI, NVIDIA, and Mistral AI — published on pages dated on or before August 14, 2026, and checked what each page actually discloses. This is an audit of disclosure, not an accusation of dishonesty. A score that only the vendor has published is not evidence of a false score, and several vendors in this dataset disclose their methodology unusually well — those rows are shown too. What the dataset measures is narrower and checkable: whether the page states the benchmark’s version, the task subset, the evaluation harness, and the reasoning-effort setting. Those four fields were checked on all 42 rows. Independent locatability — whether a third-party source reports a matching score — was attempted on only 15 of the 42 rows, and the results are labeled accordingly. The complete dataset is in this page, row by row, with the rubric stated in full so anyone can recompute the tallies. Rows that could not be resolved are reported as unresolved rather than dropped. - 01Zero of 42 rows were both fully disclosed and independently confirmed.Under the stated rubric — version, subset, harness, and effort all named on the vendor's own page, plus an independent source reporting a matching score — no row in the dataset clears both bars at once. - 02The headline finding is about disclosure, not locatability.Disclosure was checked on all 42 rows. Independent locatability was attempted on only 15 of 42 within this pass's fetch budget, so the dataset supports a disclosure rate, not a reproducibility rate. - 03The same benchmark name covers 26% and 88.4% for one model.xAI's own page reports Grok 4.6 at 26% on Terminal-Bench v3.0, while an independent evaluator separately measured the same model at 88.4% on the older v2.1 — two real numbers, two different tests, one benchmark name. - 04DeepSWE v1.1 is the dataset's strongest positive result.Four vendors' self-reported DeepSWE v1.1 scores each fell within the independent leaderboard's stated uncertainty band. Even there, every vendor page omitted at least one disclosure field — most often the harness. - 05Honest vendor-only is a real category, distinct from hidden.DeepSeek labels its internal test sets as internal; Z.ai labels its Code Bench private; Kimi K3's page is the most fully disclosed row in the set. Plain labeling tells a reader not to expect independent confirmation — which is itself disclosure. 01 — The FindingDisclosure, measured on every row. Each of the 42 rows is a vendor, model, benchmark, score triple that a vendor itself published — in a launch blog post, a model card, or a linked developer guide. For every row we read the vendor’s own page for four fields: the benchmark version where the benchmark has one , the task subset where applicable , the evaluation harness , and the reasoning-effort setting . Those four fields decide whether a reader can even identify which test produced the number, before any question of re-running it. The result: of the 41 rows we could score against the rubric — every row except Mistral’s unresolved one — only 10 name an evaluation harness, and those 10 come from just three vendors DeepSeek, Moonshot AI, and NVIDIA . Sixteen of the 41 state a reasoning-effort setting. Twenty-nine of the 42 rows are missing at least one applicable disclosure field and are scored under-specified. And zero rows in the dataset are simultaneously fully disclosed and independently confirmed — though, as the methodology below explains, the independent-confirmation column was only attempted on 15 of the 42 rows, so that zero is a statement about this dataset under this rubric, not a claim that no such row could exist. Rows across 9 vendors 41 rows we could score against the rubric plus 1 row we could not resolve at all Mistral's Shieldstral, chart-image only , kept in the table rather than dropped. Every row cites the primary page it came from. Scored rows naming a harness All 10 come from three vendors: DeepSeek a page-level footnote naming the harness and its mode , NVIDIA NeMo Gym / NeMo Evaluator SDK , and Moonshot AI's Kimi K3 page. Three further rows are harness-not-applicable. Fully specified and independently locatable No row discloses version, subset, harness, and effort on the vendor's own page and also has an independently reported matching score. The nearest misses are the four DeepSWE rows — confirmed independently, but each missing a disclosure field. Two things this finding is not. It is not a ranking of vendors by trustworthiness — a 42-row sample selected by “most recent flagship launch” says what each specific page stated, not how each company behaves in general. And it is not evidence that any number is wrong: this research re-ran no benchmark. The pattern it does document is that, at the time of writing, the standard launch-page benchmark table leaves a buyer unable to name the exact test behind most of its rows — and that the gap is disclosure the vendor could close with a footnote. 02 — MethodThe rubric, stated in full . The selection rule was fixed before any row was scored. For each of the nine vendors, we identified the single most recent flagship-model launch document dated on or before August 16, 2026, fetched it directly — never a third-party summary — and took rows from the vendor’s own comparison table in the order they appeared, skipping only competitor-model comparison columns and rows that were pricing, latency, or cost figures rather than capability scores. No row was excluded for looking strong, weak, or poorly documented. Collected: 42 vendor, model, benchmark, score rows — 41 we could score against the rubric, 1 we could not resolve at all — from 9 vendors’ own launch posts, model cards, and developer guides, each dated on or before August 14, 2026 and each fetched directly from the primary page. Checked on all 42 rows: whether the vendor’s own page states the benchmark version, task subset, evaluation harness, and reasoning-effort setting a page-level footnote covering multiple rows counts as disclosure . Attempted on 15 of 42 rows only: locating the same model, benchmark pairing on an independent leaderboard, maintainer page, or paper — the locatability column is partial, and rows not attempted are labeled “not checked this pass,” never folded into “not locatable.” Excluded: safety and red-team capability evaluations, pricing, context-window, and latency figures. Known limits: this is one bounded research pass, not a standing audit; independent leaderboards re-rank continuously; two leaderboards returned only their top rows to our fetch method; and several vendor pages render tables as chart images that yield no machine-readable numbers. As-of: vendor rows are current as of each page’s own publish date, the most recent of which is August 14, 2026; independent checks reflect each leaderboard’s state at the time of writing. Each row then receives one of four verdicts. The rubric’s most important design decision: independent reproducibility and complete disclosure are two different axes , and a row can score well on one and poorly on the other. A row that an independent leaderboard confirms can still be under-specified if the vendor’s own page never told a reader which version or harness produced the number. Fully specified, locatable Version if the benchmark has one , subset if applicable , harness, and effort are all disclosed on the vendor's own page, AND an independent source reports a matching or closely matching score for the same model. Specified but vendor-only Disclosure is complete, but no independent match was found — including benchmarks the vendor explicitly labels internal or private, where independent confirmation is not applicable by design. Under-specified At least one applicable disclosure field is missing, ambiguous, or unreachable — regardless of whether an independent match exists. All four independently confirmed DeepSWE rows land here. Not locatable The page gives only a relative claim with no absolute number to audit, or the score could not be extracted from the fetched page at all chart-image-only rendering . The independent column distinguishes five outcomes: a match within the independent source’s own stated uncertainty band, or within roughly 2 points where no band is given , a near-match a larger but small gap, stated per row , inconclusive an independent source exists and loaded, but our fetch did not surface the specific model’s row — a limit of our method, not proof the score is absent , absent checked and not found , and not checked this pass genuinely not attempted . Keeping those apart matters: collapsing “we did not check” into “nobody can find it” would be exactly the kind of imprecision this audit exists to flag. For the companion protocol on reading a single vendor table, see the four levers that decide what a vendor benchmark table means /blog/vendor-benchmark-tables-reading-disclosed-losses-2026 ; this post applies that protocol as a dataset across nine vendors. 03 — The DataThe complete 42-row dataset. Every row below cites the vendor page it came from; the group header for each vendor links the primary source. “Page footnote” means the field is disclosed in a single page-level footnote covering the vendor’s benchmark rows. “Via ‘High’ label” means the effort setting is implied by the vendor’s own column header naming the reasoning-effort SKU. Verdicts use the four-tier rubric above. | Benchmark | Score vendor page | Version | Subset | Harness | Effort | Independent check | Verdict | |---|---|---|---|---|---|---|---| | DeepSeek — V4-Pro-0813 · | launch post https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ · version labels present; no harness or effort named anywhere on the page launch post https://x.ai/news/grok-4-6 · effort implied by the “Grok 4.6 High” column label; no harness named for its own scores developer guide https://docs.z.ai/guides/llm/glm-5.3 · three tables of different disclosure character; no harness named in any of them launch announcement https://www.anthropic.com/news/claude-opus-5 · every headline evaluation is a relative claim; the page states no absolute benchmark scores GA announcement https://openai.com/index/gpt-5-6/ · per-cell version labels, no harness column, partial prose-only effort disclosure technical blog https://www.kimi.ai/blog/kimi-k3 · page-level effort disclosure plus per-benchmark footnotes naming which harness evaluated which model Hugging Face model card https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 · harness named explicitly “NeMo Gym / NeMo Evaluator SDK” ; thinking-toggle setting undisclosed announcement https://mistral.ai/news/shieldstral/ and arXiv paper https://arxiv.org/abs/2607.25857 · comparison figures rendered as chart images on both pages Two dataset notes. The release dates used for DeepSeek V4-Pro-0813 August 13 and GLM-5.3 August 14 come from our earlier reporting, not from a date field on the primary pages — the Hugging Face card and the Z.ai developer guide show no publish date. And the rows marked inconclusive against Artificial Analysis pages are inconclusive because of our fetch method: the Terminal-Bench v2.1 page states it lists 30 of 201 models, but only its top three rows loaded for us, and the Intelligence Index page surfaced only Claude Opus 5 rows. Those scores may well be present on the live pages — we simply could not confirm them, and say so rather than counting them as absent. Verdict distribution · 42 rows under the stated rubric Source: this audit's 42-row dataset, tallied from the table above 04 — Version TrapsSame name, different test. Three patterns in the dataset show why a benchmark name without its qualifiers is close to meaningless — the first two confirmed at the benchmark maintainer’s own page, the third inferred by comparing two vendors’ footnotes. Terminal-Bench v2.1 versus v3.0 xAI’s launch page reports Grok 4.6 at 26% on Terminal-Bench v3.0 — a genuinely distinct, newer track confirmed to exist at the maintainer’s leaderboard hub https://www.tbench.ai/leaderboard — while Artificial Analysis’s independent v2.1 evaluation separately measured the same model at 88.4% at “high” effort. Read side by side without the version label, those two numbers look like two different products. GLM-5.3 also reports on v3.0 28.3 , roughly in step with Grok’s v3.0 figure, while two of the three v2.1 self-reports in this dataset sit in the high 80s — DeepSeek at 87.9 and GPT-5.6 Sol at 88.8%. The third v2.1 self-report, NVIDIA’s Nemotron 3.5 Lightning at 24.58, sits down with the v3.0 figures instead, which is the limit of the version reading: the label tells a reader which test ran, not on its own how high the number will be. The maintainer’s hub currently lists five distinct tracks — 3.0, 2.1, 2.0, a legacy 1.0, and a Science track marked coming soon — so “Terminal-Bench” with no version number is presently ambiguous among at least four active or near-active tests. FrontierCode’s nested subsets The maintainer, Cognition https://cognition.com/blog/frontier-code , states the structure plainly: Diamond is the 50 hardest tasks, Main is the 100 hardest including Diamond , and Extended is the full set of 150. Extended is therefore, by construction, the easiest-skewing subset — it includes the 50 tasks the harder subsets exclude. Grok 4.6’s 61.3% is on Extended; Gemini 3.7 Flash’s 43.6% is on Main. Both pages disclose their subset, which is to their credit — but the two numbers are not the same test, and a comparison table that lines them up as if they were would mislead without either vendor having misstated anything. An “AutomationBench” name collision DeepSeek and Google both report a score on “AutomationBench” — apparently the public, GitHub-hosted tool-automation benchmark that Kimi K3’s footnotes describe evaluating on a 600-task public subset. Anthropic’s page separately reports on “Zapier AutomationBench” — by its own description a business-workflow evaluation built with Zapier specifically. These read as two different evaluations sharing a near-identical name, not three vendors citing one benchmark, and the dataset keeps them separate. An all-pass-style metric on a workflow benchmark also compresses scores in a way that can look alarming without its methodology — another reason the name alone carries so little information. None of this is new in kind — FrontierMath v2 exists because the maintainer corrected errors in a large share of the original problem set /blog/epoch-frontiermath-v2-error-corrected-ai-benchmark-analysis , and scores moved when it did. Versioning is healthy benchmark hygiene. The trap is only in citing a versioned benchmark without its version. 05 — Independent ChecksWhere the 15 attempted checks landed. Independent locatability was attempted on 15 of the 42 rows. Seven reached a conclusion: four exact matches and one near-match all on DeepSWE v1.1 and Terminal-Bench 2.1 , one confirmed absence that is a timing artifact, and one access-blocked row. Eight were inconclusive — an independent source exists and loaded, but our fetch did not surface the specific model’s row. The remaining 27 rows were either explicitly not applicable vendor-labeled private benchmarks, vendor-run benchmarks for which we found no independent leaderboard, or relative-only claims with no number to check or genuinely not attempted within this pass’s budget, and are labeled as such. The strongest positive result in the dataset is DeepSWE v1.1 — the one benchmark enough vendors reported in common to make a like-for-like check possible. The independent DeepSWE leaderboard https://deepswe.datacurve.ai/ publishes Pass@1 with an uncertainty band, average cost per task, and step count per model — a richer disclosure format than most of the vendor pages in this dataset — and four separate vendors’ self-reported scores each fell inside its band. | Model | Vendor’s own page | Independent leaderboard | Reading | |---|---|---|---| | DeepSeek V4-Pro-0813 | 62.7 | 63%±6% max | Match — inside the stated band | | Gemini 3.7 Flash | 65.3% | 65%±2% high | Match — inside the stated band | | Grok 4.6 | 65.9% | 67%±2% xhigh | Match — inside the stated band | | GPT-5.6 Sol | 72.7% | 73%±3% max | Match — inside the stated band | | GLM-5.3 | 66.9 | Absent from the snapshot | The leaderboard snapshot predates the model’s release by one day — a timing artifact | | Claude Opus 5 | Not reported on its launch page | 74%±4% max — ranked first | Leaderboard-only: the top model on the independent board never cited the benchmark itself | Even here, the disclosure gap persists: every vendor page in the match rows omitted at least one field — most often the harness — that the independent leaderboard had to supply. And the one near-match is instructive in the other direction: OpenAI’s page reports Terminal-Bench 2.1 at 88.8% without labeling which effort level produced it, while the independent evaluation lists 89.5% at “xhigh” — a 0.7-point gap that is impossible to interpret precisely because the vendor’s own effort setting is unstated. The FrontierMath row sits in the inconclusive bucket for the same methodological reason: the maintainer’s hub loaded for us without a score table, which is a statement about our fetch, not about the score. 06 — By VendorThe vendor-by-vendor scorecard . The per-vendor summary below is derived entirely from the dataset table — no new research, so every cell can be recomputed from the rows above. It is a summary of these specific rows, not a general trustworthiness ranking: each vendor is represented by one launch document, selected by recency, and a different document from the same vendor could score differently. | Vendor rows | Harness named | Effort stated | Independent checks attempted → outcome | |---|---|---|---| | DeepSeek 6 | 6 of 6 | 6 of 6 | 2 → 1 match, 1 inconclusive | | Google 5 | 0 of 5 | 0 of 5 | 2 → 1 match, 1 inconclusive | | xAI 7 | 0 of 7 | 7 of 7 — via the “High” column label | 3 → 1 match, 2 inconclusive | | Z.ai 6 | 0 of 6 | 1 of 6 | 2 → 1 absent timing artifact , 1 inconclusive | | Anthropic 4 — all relative-only claims | 0 of 4 | 1 of 4 | 0 attempted — no absolute number to check | | OpenAI 9 | 0 of 9 | 0 of 9 — partial prose disclosure only | 4 → 1 match, 1 near-match, 2 inconclusive | | Moonshot AI 1 | 1 of 1 | 1 of 1 | 0 attempted | | NVIDIA 3 | 3 of 3 | 0 of 3 | 1 → 1 inconclusive | | Mistral AI 1 — unresolved | — | — | 1 → access-blocked chart-image only | Vendors that specify their rows fully are a real result, and worth naming. DeepSeek disclosed harness and effort for all six of its rows through a single page-level footnote — “the minimal mode of DeepSeek Harness,” at max reasoning effort with temperature 1.0 and top p 0.95 — and marked its internal test sets with a dagger rather than presenting them like public benchmarks. Kimi K3’s single row is the most fully disclosed in the dataset: page-level effort settings, per-benchmark footnotes naming which of three harnesses evaluated which model, and explicit URLs to the independent sources used for competitor scores. NVIDIA named its harness on every row and its model card shows a comparison model outperforming its own model on 11 of the 14 listed benchmarks — a table that does not favor its own product on most rows shown. Anthropic’s launch page is the outlier in the other direction, in a specific and unusual way: its four headline evaluations are all relative claims — “more than doubles,” “within 0.5% of,” “three times as high as,” “around 1.5x” — with no absolute number stated for any of them, which is why all four rows land in the not-locatable tier. The counterpoint from the independent side is striking: Claude Opus 5 ranks first on the independent DeepSWE leaderboard at 74%±4%, a benchmark its own launch page never mentions. A vendor under-claiming relative to independent measurement is exactly the kind of finding a disclosure-only reading would miss. And OpenAI’s page — one of the largest benchmark tables in the dataset, with per-cell version labels — states two different Agents’ Last Exam figures on the same page, 53.6 in prose and 52.7% in its table, with no effort label connecting them. 07 — LimitsWhat this table does not show. A dataset like this invites over-reading, so the boundaries are worth stating as plainly as the findings. Not evidence that any number is false. Every vendor-only or under-specified verdict is a disclosure finding, not an accuracy finding. This research re-ran no benchmark. Not a reproducibility rate. Independent locatability was attempted on 15 of 42 rows. The other 27 are labeled not-applicable or not-checked — different claims from “not locatable,” kept apart throughout. Not a vendor trustworthiness ranking. One launch document per vendor, selected by recency, is not a representative sample of any company’s publication history. Not proof that a not-locatable score is unlocatable in general. The Claude Opus 5 ARC-AGI-3 row, for instance, points to a maintainer results page that exists and was not fetched in this pass. Not a judgment about which harness or effort setting is correct. The audit asks whether a setting was disclosed, not whether the vendor chose the right one. Three per-row caveats deserve their own line. Mistral’s WildGuardTest row stays unresolved because both the announcement and the linked arXiv abstract present the comparison as chart images — an access limitation of our text-based fetches, not a vendor disclosure failure, and the row is marked that way. OpenAI’s ARC-AGI-3 row carries a footnote marker whose text was not present in anything our fetches could load — so that row is scored as having an unreachable footnote, not as simply undisclosed. And the BrowseComp coincidence — Kimi K3 and GPT-5.6 Sol each self-report exactly 90.4 on the same benchmark under two different disclosed conditions — is reported as a coincidence a reader auditing these tables would want to find and judge for themselves. This research found no evidence either figure is wrong. and an independent source reporting a matching score. Quote the zero without the rubric and it becomes unfalsifiable — which would make it weaker, not stronger. If you cite this dataset, cite the rubric with it. 08 — For BuyersWhat to ask for before you buy. For a team using benchmark tables in model procurement, the dataset reduces to four questions to put to any vendor number — the same four fields the audit scored. Which version of the benchmark, since the dataset shows one benchmark name covering scores from the mid 20s to the high 80s? Which subset , since nested subsets like FrontierCode’s make the full set the easiest one? Which harness , since only three of nine vendors named one anywhere? And which effort setting , since a 0.7-point gap between a vendor number and an independent one cannot be interpreted without it? The realistic posture is not to distrust vendor tables but to treat them as claims pending qualification — and to prefer benchmarks with an independent leaderboard, where the check costs minutes. Our companion guide covers how to read a leaderboard without being fooled by contamination or cherry-picking /blog/llm-benchmark-methodology-2026-contamination-leaderboard-guide , and the earlier CursorBench analysis /blog/cursorbench-v3-1-vendor-benchmark-analysis explains why we found no independent leaderboard to check for that vendor-controlled benchmark — the reason those rows are marked not-applicable here rather than not-checked. This post is one of a pair of audits published together; the open-weight licence audit /blog/open-weight-model-licence-audit-2026 applies the same census method to model licences. For teams that want this class of verification run against their own shortlist before committing to a model, it is part of what our AI transformation engagements /services/ai-transformation cover. Looking forward, the incentive structure suggests disclosure will improve unevenly: benchmarks with strong independent leaderboards DeepSWE v1.1 in this dataset already show vendor self-reports converging on independently measured values, because divergence is checkable within minutes. Where no independent check exists, nothing external disciplines the footnote. If that holds, the gap this audit measures should close fastest exactly where independent infrastructure exists — which is an argument for funding leaderboards, not for distrusting vendors. 09 — ConclusionA disclosure gap a footnote could close. The gap is not honesty. It is four missing fields. Forty-two rows, nine vendors, one rubric: no row in this dataset is simultaneously fully disclosed on the vendor’s own page and independently confirmed. The finding that matters is the disclosure half — checked on every row — and it is fixable at the cost of a footnote: DeepSeek’s single page-level methodology note and Kimi K3’s per-benchmark harness footnotes show the format already exists in production. The dataset’s positive results deserve equal billing. Four vendors’ DeepSWE self-reports landed inside the independent leaderboard’s uncertainty bands. Two vendors label their private benchmarks as private, plainly. One vendor’s model tops an independent leaderboard its own launch page never cited. Disclosure and accuracy are different axes, and this dataset documents plenty of the latter where it could be checked at all — which was 15 rows out of 42. That last clause is the honest boundary of this audit, and the reason its headline is about what vendor pages state rather than what third parties can find. The table is in this page in full, with the rubric, so the next person to count can check ours.