I benchmarked what frontier models actually know about 2026 — most of it, they don't A developer built a 150-question benchmark from a 7.2M-page web index to measure frontier models' raw parametric knowledge of 2025-2026 events occurring after their training cutoffs. Across nine models, the best performer, Gemini 3.7 Flash, answered only 24.0% correctly (36/150, Wilson 95% CI [17.9%, 31.4%]), while DeepSeek-R1 followed at 20.0% and GPT-5.4 at 11.3%. A 2024-cutoff 3B control model scored 0/150 by construction, which the author cites as evidence the results reflect genuine post-cutoff signal rather than contamination. Benchmark: https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150 https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150 Dataset: https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150 https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150 Every LLM leaderboard tells you how models score on knowledge from their training data. I wanted the opposite: what do models know about things that happened AFTER their training cutoff? Not retrieval-augmented answers — raw parametric knowledge of 150 factual questions built from a 7.2M-page web index, restricted to pages citing 2025-2026 events. Two construction guarantees make contamination impossible: Scoring: case-insensitive, token-boundary containment of the gold span — a lower bound morphological variants of a correct answer can be missed . Prompts demand a bare fact, no sentence. 9 frontier models, 150 questions each, single run, no retries on the scored attempt prompt: "answer with just the requested fact, no sentence" . k = correct answers out of 150; CI = Wilson 95%. | Model | Accuracy | k/150 | Wilson 95% CI | |---|---|---|---| | Gemini 3.7 Flash | 24.0% | 36 | 17.9%, 31.4% | | DeepSeek-R1 0528 | 20.0% | 30 | 14.4%, 27.1% | | Gemini 2.5 Pro | 16.0% | 24 | 11.1%, 22.4% | | GPT-5.4 | 11.3% | 17 | 7.2%, 17.4% | | Gemini 3.5 Flash-Lite | 8.7% | 13 | 5.1%, 14.3% | | Gemini 2.5 Flash | 7.3% | 11 | 4.1%, 12.7% | | Gemma 4 31B open weights | 6.0% | 9 | 3.2%, 11.0% | | Claude Haiku 4.5 | 5.3% | 8 | 2.7%, 10.2% | | GPT-5.4 nano | 3.3% | 5 | 1.4%, 7.6% | Claude Sonnet 4.5 was scheduled three times and never started infrastructure-side, not a scoring failure — it is excluded rather than reported as zero. Gemma 4 31B hit per-question timeouts on 5 of 150 items slow open-weights serving ; the unanswered items are counted as wrong, standard practice — 9/150 = 6.0%. The best model still fails 3 questions out of 4. Gemini 3.7 Flash leads at 24% — meaning even the freshest frontier model has no parametric trace of most 2025-2026 facts. Anyone building on "the model probably knows" is wrong 76% of the time. Reasoning does not rescue knowledge. DeepSeek-R1 20.0% scores second and beats several newer generalist models — but its long chains cannot invent a fact that was never in training. Reasoning moves the needle on problems, not on missing data. Size and price do not order the ranking. GPT-5.4 11.3% sits below DeepSeek-R1 and far below Gemini 3.7 Flash; Gemini 2.5 Pro 16.0% beats its own family's newer Flash-Lite 8.7% . Freshness of training data matters more than benchmark muscle. The honest-control design works : the 2024-cutoff 3B control scores exactly 0/150 — by construction — which makes every point above zero a genuine post-cutoff signal, not contamination. Built with the kaggle-benchmarks library. Task source is public on the benchmark page — fork it and run your own lineup.