# I benchmarked what frontier models actually know about 2026 — most of it, they don't

> Source: <https://dev.to/jeffreyturov/i-benchmarked-what-frontier-models-actually-know-about-2026-most-of-it-they-dont-nm5>
> Published: 2026-10-01 14:03:21+00:00

**Benchmark:** [https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150](https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150)

**Dataset:** [https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150](https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150)

Every LLM leaderboard tells you how models score on knowledge from their training data.

I wanted the opposite: **what do models know about things that happened AFTER their training cutoff?** Not retrieval-augmented answers — raw parametric knowledge of

150 factual questions built from a 7.2M-page web index, restricted to pages citing

2025-2026 events. Two construction guarantees make contamination impossible:

Scoring: case-insensitive, token-boundary containment of the gold span — a lower

bound (morphological variants of a correct answer can be missed). Prompts demand

a bare fact, no sentence.

9 frontier models, 150 questions each, single run, no retries on the scored attempt

(prompt: "answer with just the requested fact, no sentence"). k = correct answers

out of 150; CI = Wilson 95%.

| Model | Accuracy | k/150 | Wilson 95% CI | 
|---|---|---|---|
| Gemini 3.7 Flash | **24.0%** | 36 | [17.9%, 31.4%] | 
| DeepSeek-R1 (0528) | **20.0%** | 30 | [14.4%, 27.1%] | 
| Gemini 2.5 Pro | **16.0%** | 24 | [11.1%, 22.4%] | 
| GPT-5.4 | **11.3%** | 17 | [7.2%, 17.4%] | 
| Gemini 3.5 Flash-Lite | **8.7%** | 13 | [5.1%, 14.3%] | 
| Gemini 2.5 Flash | **7.3%** | 11 | [4.1%, 12.7%] | 
| Gemma 4 31B (open weights) | **6.0%** | 9 | [3.2%, 11.0%] | 
| Claude Haiku 4.5 | **5.3%** | 8 | [2.7%, 10.2%] | 
| GPT-5.4 nano | **3.3%** | 5 | [1.4%, 7.6%] | 

Claude Sonnet 4.5 was scheduled three times and never started (infrastructure-side,

not a scoring failure) — it is excluded rather than reported as zero. Gemma 4 31B

hit per-question timeouts on 5 of 150 items (slow open-weights serving); the

unanswered items are counted as wrong, standard practice — 9/150 = 6.0%.

**The best model still fails 3 questions out of 4.** Gemini 3.7 Flash leads at

24% — meaning even the freshest frontier model has no parametric trace of most

2025-2026 facts. Anyone building on "the model probably knows" is wrong 76% of

the time.

**Reasoning does not rescue knowledge.** DeepSeek-R1 (20.0%) scores second and

beats several newer generalist models — but its long chains cannot invent a fact

that was never in training. Reasoning moves the needle on problems, not on

missing data.

**Size and price do not order the ranking.** GPT-5.4 (11.3%) sits below

DeepSeek-R1 and far below Gemini 3.7 Flash; Gemini 2.5 Pro (16.0%) beats its own

family's newer Flash-Lite (8.7%). Freshness of training data matters more than

benchmark muscle.

**The honest-control design works**: the 2024-cutoff 3B control scores exactly

0/150 — by construction — which makes every point above zero a genuine

post-cutoff signal, not contamination.

Built with the `kaggle-benchmarks` library. Task source is public on the

benchmark page — fork it and run your own lineup.
