cd /news/ai-research/my-factual-recall-tasks-were-scoring… · home topics ai-research article
[ARTICLE · art-138629] src=dev.to ↗ pub= topic=ai-research verified=true sentiment=↓ negative

My factual-recall tasks were scoring format, not facts

A developer's drift-detection harness revealed that two factual-recall tasks, fact-element and fact-planet, were grading output format rather than knowledge, with Claude Sonnet 5 failing the gold-symbol task 60 out of 60 times and GPT-4o mini failing 59 out of 59 despite both models answering correctly. The near-identical control task fact-capital, which explicitly instructs models to answer with only the city name, passed cleanly across all runs, showing the failures stem from models adding punctuation or markdown rather than lacking the facts. The developer estimates the two tasks account for 5.7 points of every affected model's score and skews the public board's factual-recall category breakdown.

by read4 min views1 publishedSep 23, 2026

Originally published at erikhill.dev. The numbers below are checked against the repository they come from.

This is a finding about my own harness. The suspect is the probe, not the models it measures.

I built a detector that decides whether a day's drift run contains anything worth

writing up. The first thing it did was accuse my own suite. One task,

fact-element, had failed on three or more providers on the same day, on 46

separate days.

The task is this:

What is the chemical symbol for gold? Two letters only.

Graded by exact match against au, after stripping whitespace and lowercasing.

Claude Sonnet 5 failed it 60 times out of 60. GPT-4o mini failed it 59 out of 59. Llama 3.1 8B, 23 out of 23. Those are not knowledge results. A model

claim how it was checked
fact-element failed on 3 or more providers on the same day, 46 times cross-provider flip analysis over every stored run in dashboard/drift_board.json
those were graded failures, not outages every run counted had reliability: 1.0 , so no call errored and none was truncated
the models know the answer gpt-4o-mini answered'The chemical symbol for gold is Au.' ;claude-sonnet-5 answered'**Au**\n\nThat's the chemical symbol for gold, derived from the Latin word *aurum*.'
the same models pass a near-identical task all three answered fact-capital with exactly'Tokyo' , 0 failures in 59 and 60 runs
one trailing period is the whole difference gpt-4o-mini answeredfact-planet with'Mercury.' and was recorded wrong 59 times out of 59
a model that formats tersely passes cleanly claude-opus-4-8 answered'Au' ,'Mercury' ,'Tokyo' , and failsfact-element on 1 of 60 runs

The answers above are not from the board. They came from nine live calls made

through the probe's own call path, printed with repr() so punctuation and

markdown are visible rather than inferred.

fact-capital is the control, and it is the reason this is a finding rather

than a theory. It asks for the capital of Japan and every model in the table

answers 'Tokyo', exactly, on every run. Same grader, same strictness, same

models, same day. Nothing about exact matching is broken in general.

The three tasks differ in one way. fact-capital says "Answer with only the city name", which is an instruction about the shape of the output.

fact-element says fact-planet says The suite has 35 tasks, so the smallest accuracy change it can print is

100/35 = 2.86 points. fact-element and fact-planet together are 5.7 points

of every affected model's score, permanently, for reasons that have nothing to

do with the models.

Worse than the points: these tasks are tagged factual-recall, and that tag

feeds a per-category breakdown on the public board. GPT-4o mini's

factual-recall score has been reading as a knowledge number when two thirds of

it is a formatting number. Anyone comparing models on that column, including

me, was comparing how terse they are.

The part I am least sure of is the explanation, not the measurement. The numbers

are solid and the live answers are quoted verbatim, but "models treat answer with only X as a constraint and

Two things genuinely argue against my reading. Grok 4.5 fails fact-element on

26 of 49 runs, which is close to a coin flip, and a rigid house style should not

produce that. And Claude Sonnet 5 failed fact-planet on 38 of 60 historical

runs but answered 'Mercury' cleanly when I called it today, so at least some

of this moves over time and is not a fixed property of the model.

There is also a real question about what the fix should be, and I do not think

it is obviously "loosen the grader". A grader that accepts 'Mercury.' also

accepts a model that ignores the instruction, and the instruction was part of

the task. Making the grader lenient would convert a visible measurement problem

into an invisible one. The alternative is to say plainly that these are

instruction-following tasks and retag them, which changes what the board has

been reporting for three months.

If you maintain an eval suite with exact-match graders on short answers: check what your passing models actually return, not just whether they passed. The

board I built recorded which tasks failed for three months and never once

recorded what the model said, which is why this took a live call to see. That

gap is the more embarrassing half of this post.

I build deterministic evaluation and verification tooling for LLM systems, and I am looking for my first full-time role in AI evaluation or QA engineering. Remote US Eastern, or Charleston SC. https://erikhill.dev

Live board: [https://egnaro9.github.io/model-drift/](https://egnaro9.github.io/model-drift/)

Suite, graders and runner: [https://github.com/egnaro9/model-drift](https://github.com/egnaro9/model-drift)

The finding was surfaced automatically by the drift detector. The investigation, the live calls and the judgement of whether it was worth saying are mine.

── more in #ai-research 4 stories · sorted by recency
── more on @claude sonnet 5 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/my-factual-recall-ta…] indexed:0 read:4min 2026-09-23 ·