cd /news/large-language-models/i-benchmarked-what-frontier-models-a… · home › topics › large-language-models › article
[ARTICLE · art-143242] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=↓ negative

I benchmarked what frontier models actually know about 2026 — most of it, they don't

A developer built a 150-question benchmark from a 7.2M-page web index to measure frontier models' raw parametric knowledge of 2025-2026 events occurring after their training cutoffs. Across nine models, the best performer, Gemini 3.7 Flash, answered only 24.0% correctly (36/150, Wilson 95% CI [17.9%, 31.4%]), while DeepSeek-R1 followed at 20.0% and GPT-5.4 at 11.3%. A 2024-cutoff 3B control model scored 0/150 by construction, which the author cites as evidence the results reflect genuine post-cutoff signal rather than contamination.

by read2 min views4 publishedOct 1, 2026
**Benchmark:** [https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150](https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150)

**Dataset:** [https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150](https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150)

Every LLM leaderboard tells you how models score on knowledge from their training data.

I wanted the opposite: what do models know about things that happened AFTER their training cutoff? Not retrieval-augmented answers — raw parametric knowledge of

150 factual questions built from a 7.2M-page web index, restricted to pages citing

2025-2026 events. Two construction guarantees make contamination impossible:

Scoring: case-insensitive, token-boundary containment of the gold span — a lower

bound (morphological variants of a correct answer can be missed). Prompts demand

a bare fact, no sentence.

9 frontier models, 150 questions each, single run, no retries on the scored attempt

(prompt: "answer with just the requested fact, no sentence"). k = correct answers

out of 150; CI = Wilson 95%.

Model Accuracy k/150 Wilson 95% CI
Gemini 3.7 Flash 24.0% 36 [17.9%, 31.4%]
DeepSeek-R1 (0528) 20.0% 30 [14.4%, 27.1%]
Gemini 2.5 Pro 16.0% 24 [11.1%, 22.4%]
| GPT-5.4 | **11.3%** | 17 | [7.2%, 17.4%] | 
| Gemini 3.5 Flash-Lite | **8.7%** | 13 | [5.1%, 14.3%] | 

| Gemini 2.5 Flash | 7.3% | 11 | [4.1%, 12.7%] | | Gemma 4 31B (open weights) | 6.0% | 9 | [3.2%, 11.0%] | | Claude Haiku 4.5 | 5.3% | 8 | [2.7%, 10.2%] |

| GPT-5.4 nano | 3.3% | 5 | [1.4%, 7.6%] | Claude Sonnet 4.5 was scheduled three times and never started (infrastructure-side,

not a scoring failure) — it is excluded rather than reported as zero. Gemma 4 31B

hit per-question timeouts on 5 of 150 items (slow open-weights serving); the unanswered items are counted as wrong, standard practice — 9/150 = 6.0%.

The best model still fails 3 questions out of 4. Gemini 3.7 Flash leads at

24% — meaning even the freshest frontier model has no parametric trace of most

2025-2026 facts. Anyone building on "the model probably knows" is wrong 76% of

the time.

Reasoning does not rescue knowledge. DeepSeek-R1 (20.0%) scores second and

beats several newer generalist models — but its long chains cannot invent a fact

that was never in training. Reasoning moves the needle on problems, not on

missing data.

Size and price do not order the ranking. GPT-5.4 (11.3%) sits below

DeepSeek-R1 and far below Gemini 3.7 Flash; Gemini 2.5 Pro (16.0%) beats its own

family's newer Flash-Lite (8.7%). Freshness of training data matters more than

benchmark muscle.

The honest-control design works: the 2024-cutoff 3B control scores exactly

0/150 — by construction — which makes every point above zero a genuine

post-cutoff signal, not contamination.

Built with the kaggle-benchmarks library. Task source is public on the

benchmark page — fork it and run your own lineup.

── more in #large-language-models 4 stories · sorted by recency
── more on @gemini 3.7 flash 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-benchmarked-what-f…] indexed:0 read:2min 2026-10-01 · —