# Does your model know when it doesn't know? A benchmark for the ESCALATE answer

> Source: <https://dev.to/sean_campbell_840bd62bf7e/does-your-model-know-when-it-doesnt-know-a-benchmark-for-the-escalate-answer-268o>
> Published: 2026-09-30 09:36:03+00:00

*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*

Most leaderboards ask one question: did the model get it right? I wanted to ask a second one: **does the model know when it can't?**

I run a small multi-agent system on one laptop, where local models hand work up to bigger ones. In a setup like that, a small model that answers wrong is worse than one that says "I can't do this, pass it up." A modest model that knows its limits can sit safely in the chain. A confident one that bluffs can't.

So every task in this benchmark has a refusal token, `ESCALATE`. In **one item out of five, the answer has been deliberately removed**, or the document doesn't contain it. On those items, `ESCALATE` is the only correct reply.

There are 200 items across four everyday job shapes:

Every model gets **two scores**: its *task score* on the answerable items, and its *false-confidence rate*, meaning how often it answered anyway when the right reply was `ESCALATE`. Each answer also carries a stated confidence, so I can draw a reliability diagram too.

Every item is invented from scratch, and nothing is scraped. A privacy gate checks the whole set before it's published.

Two groups, on one chart:

The question behind the chart: *can a laptop's 3B know its own limits as well as a frontier model knows its own?*

*The runs are in progress. Before any model touched the fixtures, I wrote down my predictions and timestamped them so they can't drift toward the results:*

I'll grade those here, misses included, once the numbers are in.

*Kaggle link coming once the benchmark is published there.*
