This is a submission for the Kaggle Benchmarking Challenge Most leaderboards ask one question: did the model get it right? I wanted to ask a second one: does the model know when it can't?
I run a small multi-agent system on one laptop, where local models hand work up to bigger ones. In a setup like that, a small model that answers wrong is worse than one that says "I can't do this, pass it up." A modest model that knows its limits can sit safely in the chain. A confident one that bluffs can't.
So every task in this benchmark has a refusal token, ESCALATE. In one item out of five, the answer has been deliberately removed, or the document doesn't contain it. On those items, ESCALATE is the only correct reply.
There are 200 items across four everyday job shapes:
Every model gets two scores: its task score on the answerable items, and its false-confidence rate, meaning how often it answered anyway when the right reply was ESCALATE. Each answer also carries a stated confidence, so I can draw a reliability diagram too.
Every item is invented from scratch, and nothing is scraped. A privacy gate checks the whole set before it's published.
Two groups, on one chart:
The question behind the chart: can a laptop's 3B know its own limits as well as a frontier model knows its own?
The runs are in progress. Before any model touched the fixtures, I wrote down my predictions and timestamped them so they can't drift toward the results:
I'll grade those here, misses included, once the numbers are in.
Kaggle link coming once the benchmark is published there.