**Benchmark:** [https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150](https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150)
**Dataset:** [https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150](https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150)
Every LLM leaderboard tells you how models score on knowledge from their training data.
I wanted the opposite: what do models know about things that happened AFTER their training cutoff? Not retrieval-augmented answers — raw parametric knowledge of
150 factual questions built from a 7.2M-page web index, restricted to pages citing
2025-2026 events. Two construction guarantees make contamination impossible:
Scoring: case-insensitive, token-boundary containment of the gold span — a lower
bound (morphological variants of a correct answer can be missed). Prompts demand
a bare fact, no sentence.
9 frontier models, 150 questions each, single run, no retries on the scored attempt
(prompt: "answer with just the requested fact, no sentence"). k = correct answers
out of 150; CI = Wilson 95%.
| Model | Accuracy | k/150 | Wilson 95% CI |
|---|---|---|---|
| Gemini 3.7 Flash | 24.0% | 36 | [17.9%, 31.4%] |
| DeepSeek-R1 (0528) | 20.0% | 30 | [14.4%, 27.1%] |
| Gemini 2.5 Pro | 16.0% | 24 | [11.1%, 22.4%] |
| GPT-5.4 | **11.3%** | 17 | [7.2%, 17.4%] |
| Gemini 3.5 Flash-Lite | **8.7%** | 13 | [5.1%, 14.3%] |
| Gemini 2.5 Flash | 7.3% | 11 | [4.1%, 12.7%] | | Gemma 4 31B (open weights) | 6.0% | 9 | [3.2%, 11.0%] | | Claude Haiku 4.5 | 5.3% | 8 | [2.7%, 10.2%] |
| GPT-5.4 nano | 3.3% | 5 | [1.4%, 7.6%] | Claude Sonnet 4.5 was scheduled three times and never started (infrastructure-side,
not a scoring failure) — it is excluded rather than reported as zero. Gemma 4 31B
hit per-question timeouts on 5 of 150 items (slow open-weights serving); the unanswered items are counted as wrong, standard practice — 9/150 = 6.0%.
The best model still fails 3 questions out of 4. Gemini 3.7 Flash leads at
24% — meaning even the freshest frontier model has no parametric trace of most
2025-2026 facts. Anyone building on "the model probably knows" is wrong 76% of
the time.
Reasoning does not rescue knowledge. DeepSeek-R1 (20.0%) scores second and
beats several newer generalist models — but its long chains cannot invent a fact
that was never in training. Reasoning moves the needle on problems, not on
missing data.
Size and price do not order the ranking. GPT-5.4 (11.3%) sits below
DeepSeek-R1 and far below Gemini 3.7 Flash; Gemini 2.5 Pro (16.0%) beats its own
family's newer Flash-Lite (8.7%). Freshness of training data matters more than
benchmark muscle.
The honest-control design works: the 2024-cutoff 3B control scores exactly
0/150 — by construction — which makes every point above zero a genuine
post-cutoff signal, not contamination.
Built with the kaggle-benchmarks library. Task source is public on the
benchmark page — fork it and run your own lineup.