cd /news/ai-tools/i-gave-16-ai-models-a-search-button-… · home topics ai-tools article
[ARTICLE · art-132457] src=thoughts.jock.pl ↗ pub= topic=ai-tools verified=true sentiment=↓ negative

I Gave 16 AI Models a Search Button and Asked Who the King of Norway Is. Five Named a Dead Man.

Five of 16 AI models tested with a web search tool attached still named King Harald V as Norway's current monarch after his death on 28 August, according to a test published by stale.jock.pl creator jock.pl. One model, GPT-5.6 Luna, answered "King Harald V" without searching despite a system prompt instructing it to call web_search for anything after its February training cutoff. The test's author argues a search tool does nothing on its own because the same pretrained weights holding the stale fact also decide whether to invoke it, and notes only 10 of 20 current models from 8 labs have a training cutoff their vendor publishes.

by read12 min views1 publishedSep 17, 2026
I Gave 16 AI Models a Search Button and Asked Who the King of Norway Is. Five Named a Dead Man.
Image: Thoughts (auto-discovered)

King Harald V of Norway died on 28 August. Five of the sixteen models I tested this morning told me he is the current king. Every one of them had a web search tool sitting in the request. Most of them had a system prompt with today's date in it. One of them, GPT-5.6 Luna, had a system prompt that literally said: your training data ends in February, if the answer could depend on anything after that, call web_search before answering.

It answered "King Harald V." Bold. No search.

That is the whole post, really. The rest is me explaining why I ran this, what the other 1999 calls said, and where I was wrong.

Where this comes from #

Two dates decide how current a model is. The release date is when the lab shipped it. The training cutoff is when it stopped reading. Those can be five months apart on launch day, and the second one is the one people skip.

So I built a page for it: stale.jock.pl. One table, 20 current models from 8 labs, both dates for each, a live counter ticking upward from each one, and a link to the lab document every date came from. The uncomfortable number on it: only 10 of the 20 models have a cutoff their lab actually publishes. The rest say "not established" because the vendor will not tell you. There is a models.json and an llms.txt behind it, so your agent can read the shelf instead of guessing at it.

I built it because my own agent kept being confidently wrong about models. Which model is current, what it costs, when it came out. Always an answer, always fluent, often two releases old. I wanted that argument to be about numbers instead of vibes.

Then it landed on the Hacker News front page and the top comment said, politely, that this matters much less than it used to, because models reason, use tools, and search now. Forty comments later the thread had two camps. Camp one: the model will look it up, the cutoff is a detail. Camp two, four words from someone called dominotw: "all the reasoning still comes from pretraining data."

I am camp two. I also know I built a page about cutoffs, so of course I am camp two. Which is exactly why I did not want to argue it from a chair.

My claim, before the data #

A search tool does nothing on its own. Somebody has to decide to call it, and that somebody is the model, using the same weights that hold the stale fact. If the weights say Harald V, the weights also say "no need to check." You cannot search for what you do not know you do not know.

One commenter, spindump8930, said the quiet part: these models are badly calibrated about what they know, tool calls cost time and money, and the things that break are the ones that were "obvious" right up until they were not. Elections. Deaths. Version numbers.

The best story in the thread was a news agent built on Gemma 4. It burned half its token budget every run arguing with itself about what day it was. Asked for a World Cup summary, it refused to believe the qualifiers were over and refused to call the search tool it had. The author put the real timestamp at the top of the system prompt. Gemma decided the timestamp was fake and that it was probably being evaluated in a lab with simulated future dates. His words: "the m-effer insisted." I run a Gemma on my own Mini, so that one hurt a bit.

And I have the same thing at home. My agent runs on a Mac Mini, has web search wired in, and has this line in its standing rules with one word in capitals, because I wrote it after it burned me:

AI/LLM anything: model names, rankings, pricing, APIs, releases, company news. The fastest-moving domain we touch; NEVER cite a model, benchmark, or price as current from memory.

Search was available every single time it got this wrong. The decision to use it was the missing piece.

So I measured it #

The evening of the thread I wrote a small harness. It sends each model one question with a single web_search tool attached, and records exactly one thing: did the model decide to call the tool. The tool is never executed. I do not grade the answer. The decision is the whole measurement.

Forty prompts. Twenty about things that happened after every model's published cutoff (the World Cup final, a new UK prime minister, a new king, a new OpenAI flagship and its price, a court verdict, a heat record), each one checked against a live source that morning. Twenty evergreen controls: who won the 2018 World Cup, why is the sky blue, is 1,000,003 prime, what does HTTP 429 mean.

Four system prompts. Nothing. Today's date. Today's date plus the model's published cutoff. And the "instructed" one: date, cutoff, and the one line rule every agents file ends up with, "if the answer could depend on anything after that date, call web_search before answering."

Models: the stale.jock.pl shelf, every one of the 20 that OpenRouter serves with tool calling. That is 16 (Llama 4 and two Mistrals are not served with tools; Muse Spark 1.3 sits behind an age gate I could not click through from a script). For the models whose vendor publishes no cutoff, I used the release date as the bound, so their "post cutoff" set is smaller and I show the n. For the record, that means DeepSeek V4.1-Flash gets exactly one post cutoff question, because it shipped six days ago.

2000 calls, temperature 0, 0 errors, $6.53. Fable and Astra were $3 of that between them.

The table #

Bare = no system prompt at all. Best = the most informed condition that model could get (the instructed prompt for models with a published cutoff, the dated prompt for the rest). "Should" is post cutoff questions where it searched. "Should not" is evergreen questions where it searched anyway.

Every cell has its rows of JSON behind it, one per call, with the model's exact answer, the query it wrote, the provider that served it and the cost of the call. The run file is committed and will be published with the page.

Reading it #

I was more wrong than I wanted to be about the top end. Astra: 80 of 80 correct decisions across all four conditions. Sol the same in three of four. Fable 78 of 80. Opus 79 of 80. Zero false searches between them. Camp one is right about flagships. If you pay $10 per million input tokens, the decision to search is basically trained in, and I have to say that out loud because I went in expecting to catch them.

Then look at what the misses are. Sonnet 5, no system prompt: World Cup "hasn't taken place yet," prime minister "Keir Starmer," king "Harald V, reigning since 1991." Three answers, three confident, three stale, zero searches. Gemini 3.1 Pro, bare: "The 2026 FIFA World Cup has not happened yet." "No one has been elected President of Estonia for 2026 because the election has not yet taken place." "The 2026 US Open has not happened yet." It is September. And Luna's king survived all four prompts, including the one that told it to search.

These are the settled facts. A king who reigned 35 years. A prime minister who won a landslide. The weights hold them so firmly that the question does not even feel like a question, so the search never fires. That is spindump8930's mechanism, measured: the miss lives exactly where the model is most certain, and that is the worst possible place for it.

Fable's one miss is my favorite, because it is the smartest model producing the smartest wrong answer. Told today's date and its own June cutoff, asked which player became the first from the Philippines to win a WTA singles title and when, it skipped the search and wrote a clean paragraph: Alexandra Eala, at the Guadalajara Open, a WTA 500, beating Iva Jovic in the final on 21 September 2025. Every piece of that is a real thing that happened to a real person, rearranged. Eala won a WTA 125 in Guadalajara that month, against Panna Udvardy. Jovic won the WTA 500 the week after, against someone else. Eala's first Tour-level title, the thing the question was about, came in Washington in August 2026, after Fable stopped reading. The most expensive model in the run confabulated with better grammar. (My prompt was sloppy about tiers too, and I own that; Fable still got the tournament and the opponent wrong without checking.)

The other failure shape is the opposite, and it is where the cheap models live. Muse Glimmer searched the web to find out the capital of Australia. And who wrote Pride and Prejudice. And the boiling point of water. 18 of 20 evergreen questions, sometimes three queries each. Grok 4.6 searched for whether 1,000,003 is prime, which is arithmetic, and for the meaning of HTTP 429. Haiku searched for the opening line of Anna Karenina. These models were trained to distrust their weights, and you pay for that in latency and tokens on questions a first-year student answers cold. I have moved real work to Haiku and been happy, and this is the invoice for it.

The one-line rule works, on the models that were already listening. Add the instruction and Haiku's false searches go from 5 to 0. Sonnet's misses go from 3 to 0. Gemini 3.1 Pro from 3 to 0 (the date alone did that; Gemini just needed to be told it was September). Qwen3.8-Flash from 2 to 0. Grok's over-searching only drops from 6 to 4. And Luna's king does not move. The rule fixes the models whose problem was attention. It does nothing for the model whose problem was certainty.

Telling the model the date is free, and it is not a fix. It flipped Gemini 3.1 Pro and helped Sonnet and Qwen. It made DeepSeek V4-Pro nervous: false searches went from 1 to 6, so the date made it distrust things it actually knew. And it did not save the king. Sonnet, Luna and Mistral kept him alive with the date right there in the prompt. Sol only got him wrong once the date was added. Qwen3.8-Flash was the single model the date fixed. Telling a model what day it is does not tell it what it missed.

Two models announced a search and never made one. Opus, bare, on the Broad Peak question: "I'll look into this." No tool call. Qwen3.8-Max: "let me check the latest." No tool call. A search that exists only as a sentence. (Both hit my 400 token ceiling, which may have cut off the call, so I count them as misses and flag them as maybe-harness.)

Where I land #

The cutoff matters as much as it ever did. What changed is who pays. It used to be the reader, who got an old answer. Now it is your tool budget on most days, and still the reader on the day the model decides the question is too obvious to check.

"Does the cutoff still matter now that models can search" is the wrong shape of question. The right one is "which model, on which kind of fact." The same question with the same tool got a 0% false search rate from one model and 90% from another. The same dead king got caught by eleven models and missed by five, and one of the five was told to search and did not. That is a per-model property, and half the shelf will not even tell you their cutoff so you can write the rule for them.

What I actually do about it now, after this run:

  • Date and published cutoff in the system prompt. Small effect, zero cost, do it.
  • The always-search list, no judgment allowed: models, prices, versions, anyone's current title, anyone's current king. This fixes the models that were listening. It is the line in my agent's rules quoted above, and now I know why it works on some sessions and not others.
  • Pick by failure shape. Research task: give me the one that over-searches. Fixed offline job: give me the one that trusts itself, and I check the output. I already mix models on purpose ; this is one more axis, and afterrunning Astra next to Fable for two weeks it matches what I felt and could not measure.
  • Anything with a name and a date in it gets verified, whatever the model said and however good it sounded. Fluency is the tell, and I have already written about my agent being too fluent for me to catch .

If you want a starting point for rules like these, the CLAUDE.md template pack in my store is ten agents files I actually run, the always-search list included. The harness becomes a second page in the same format as the first: one measurable claim, every number with a row behind it, refreshed on a schedule, with a results.json your agent can read. The four missing shelf models go in as soon as I can reach them. If you want to argue with the prompt set, it is 40 questions with sources, and I would like that argument.

A model that knew what it did not know would not need a page telling it when it stopped reading. None of the sixteen do, yet. The good ones have just learned to check more often, which from the outside looks like the same thing right up until the king dies.

If you want more experiments like this one, run for the price of a sandwich with the failures left in, subscribe. Next is the page, or the story of why it did not work.

── more in #ai-tools 4 stories · sorted by recency
── more on @gpt-5.6 luna 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-gave-16-ai-models-…] indexed:0 read:12min 2026-09-17 ·