{"slug": "i-benchmarked-what-frontier-models-actually-know-about-2026-most-of-it-they-don", "title": "I benchmarked what frontier models actually know about 2026 — most of it, they don't", "summary": "A developer built a 150-question benchmark from a 7.2M-page web index to measure frontier models' raw parametric knowledge of 2025-2026 events occurring after their training cutoffs. Across nine models, the best performer, Gemini 3.7 Flash, answered only 24.0% correctly (36/150, Wilson 95% CI [17.9%, 31.4%]), while DeepSeek-R1 followed at 20.0% and GPT-5.4 at 11.3%. A 2024-cutoff 3B control model scored 0/150 by construction, which the author cites as evidence the results reflect genuine post-cutoff signal rather than contamination.", "body_md": "**Benchmark:** [https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150](https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150)\n\n**Dataset:** [https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150](https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150)\n\nEvery LLM leaderboard tells you how models score on knowledge from their training data.\n\nI wanted the opposite: **what do models know about things that happened AFTER their training cutoff?** Not retrieval-augmented answers — raw parametric knowledge of\n\n150 factual questions built from a 7.2M-page web index, restricted to pages citing\n\n2025-2026 events. Two construction guarantees make contamination impossible:\n\nScoring: case-insensitive, token-boundary containment of the gold span — a lower\n\nbound (morphological variants of a correct answer can be missed). Prompts demand\n\na bare fact, no sentence.\n\n9 frontier models, 150 questions each, single run, no retries on the scored attempt\n\n(prompt: \"answer with just the requested fact, no sentence\"). k = correct answers\n\nout of 150; CI = Wilson 95%.\n\n| Model | Accuracy | k/150 | Wilson 95% CI | \n|---|---|---|---|\n| Gemini 3.7 Flash | **24.0%** | 36 | [17.9%, 31.4%] | \n| DeepSeek-R1 (0528) | **20.0%** | 30 | [14.4%, 27.1%] | \n| Gemini 2.5 Pro | **16.0%** | 24 | [11.1%, 22.4%] | \n| GPT-5.4 | **11.3%** | 17 | [7.2%, 17.4%] | \n| Gemini 3.5 Flash-Lite | **8.7%** | 13 | [5.1%, 14.3%] | \n| Gemini 2.5 Flash | **7.3%** | 11 | [4.1%, 12.7%] | \n| Gemma 4 31B (open weights) | **6.0%** | 9 | [3.2%, 11.0%] | \n| Claude Haiku 4.5 | **5.3%** | 8 | [2.7%, 10.2%] | \n| GPT-5.4 nano | **3.3%** | 5 | [1.4%, 7.6%] | \n\nClaude Sonnet 4.5 was scheduled three times and never started (infrastructure-side,\n\nnot a scoring failure) — it is excluded rather than reported as zero. Gemma 4 31B\n\nhit per-question timeouts on 5 of 150 items (slow open-weights serving); the\n\nunanswered items are counted as wrong, standard practice — 9/150 = 6.0%.\n\n**The best model still fails 3 questions out of 4.** Gemini 3.7 Flash leads at\n\n24% — meaning even the freshest frontier model has no parametric trace of most\n\n2025-2026 facts. Anyone building on \"the model probably knows\" is wrong 76% of\n\nthe time.\n\n**Reasoning does not rescue knowledge.** DeepSeek-R1 (20.0%) scores second and\n\nbeats several newer generalist models — but its long chains cannot invent a fact\n\nthat was never in training. Reasoning moves the needle on problems, not on\n\nmissing data.\n\n**Size and price do not order the ranking.** GPT-5.4 (11.3%) sits below\n\nDeepSeek-R1 and far below Gemini 3.7 Flash; Gemini 2.5 Pro (16.0%) beats its own\n\nfamily's newer Flash-Lite (8.7%). Freshness of training data matters more than\n\nbenchmark muscle.\n\n**The honest-control design works**: the 2024-cutoff 3B control scores exactly\n\n0/150 — by construction — which makes every point above zero a genuine\n\npost-cutoff signal, not contamination.\n\nBuilt with the `kaggle-benchmarks` library. Task source is public on the\n\nbenchmark page — fork it and run your own lineup.", "url": "https://wpnews.pro/news/i-benchmarked-what-frontier-models-actually-know-about-2026-most-of-it-they-don", "canonical_source": "https://dev.to/jeffreyturov/i-benchmarked-what-frontier-models-actually-know-about-2026-most-of-it-they-dont-nm5", "published_at": "2026-10-01 14:03:21+00:00", "updated_at": "2026-10-01 14:14:42.074946+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "machine-learning", "artificial-intelligence"], "entities": ["Gemini 3.7 Flash", "DeepSeek-R1", "Gemini 2.5 Pro", "GPT-5.4", "Gemma 4 31B", "Claude Haiku 4.5", "Kaggle", "Claude Sonnet 4.5"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-benchmarked-what-frontier-models-actually-know-about-2026-most-of-it-they-don", "markdown": "https://wpnews.pro/news/i-benchmarked-what-frontier-models-actually-know-about-2026-most-of-it-they-don.md", "text": "https://wpnews.pro/news/i-benchmarked-what-frontier-models-actually-know-about-2026-most-of-it-they-don.txt", "jsonld": "https://wpnews.pro/news/i-benchmarked-what-frontier-models-actually-know-about-2026-most-of-it-they-don.jsonld"}}