# Retrieval confidence can't tell your RAG chatbot when the answer is missing

> Source: <https://dev.to/klausbyskov/retrieval-confidence-cant-tell-your-rag-chatbot-when-the-answer-is-missing-2ml0>
> Published: 2026-10-08 07:38:21+00:00

*I ran 65 questions against our own knowledge base and checked which ones the retrieved text actually answered. The similarity scores overlapped too much for any threshold to work.*

Our own handoff docs used to describe it: if retrieval confidence falls below a threshold, hand the visitor to a human instead of generating a weak answer. The default was 0.4. It's the design most people building a RAG chatbot reach for, and I no longer think it works. This post is about why: first a hunch from production, then a small experiment.

In a typical RAG pipeline, including ours at Asktopus, the visitor's question is embedded (we use OpenAI's text-embedding-3-small). The vector database (Qdrant) returns the five closest chunks by cosine similarity, and those chunks go into the prompt. The best of those five scores is what usually gets called "retrieval confidence".

It's free, already computed and between 0 and 1, so it looks like the system telling you how sure it is.

```
if confidence < THRESHOLD:
    hand_off_to_human()      # or reply "I don't know"
else:
    answer_from_context()
```

It promises two things at once: no made-up answers when the knowledge is thin, and a human for the questions the bot can't handle. All you have to do is pick the number.

In production, the scores of 101 real answers on one site ranged from 0.11 to 0.67, with a median of 0.36. At the 0.4 our docs recommended, more than half of them would have gone to a human. I couldn't tell from the numbers which of those answers were good, and 101 answers on one site is far too few to tune a threshold on. So I set up a test where I could check.

The knowledge base is asktopus.com itself: 44 crawled pages in English and Danish, split into 219 chunks of up to 1,500 characters. The pipeline is exactly production's: same embedding model, same Qdrant search, top 5, confidence = the best score.

I wrote 65 questions: 40 I expected the site to answer, 20 I expected it not to (a Zapier integration, ISO 27001, an uptime SLA), and 5 plainly off-topic (the capital of France, tomorrow's weather). Then I set my expectations aside and labelled each question by reading the five chunks the model would receive. Does this text contain the answer? **Yes**, **partial** (it answers a narrower version: conversations can be exported, but the format isn't stated), or **no**.

My expectations were off. Of the 40 questions I expected the site to answer, only 24 got retrieved text that answered them; 6 were partial and 10 got nothing useful. Three of those ten were retrieval misses: the answer was in the knowledge base, just not in the top five. All 25 questions I expected it not to answer got a no.

| Label | Questions | Min | Median | Max | 
|---|---|---|---|---|
| Yes | 24 | 0.309 | 0.499 | 0.708 | 
| Partial | 6 | 0.331 | 0.379 | 0.582 | 
| No | 35 | 0.102 | 0.345 | 0.590 | 

The score isn't noise: answerable questions score higher on average. But the two groups overlap from 0.31 to 0.59, and that band holds 19 of the 24 answerable questions, 23 of the 35 unanswerable ones and all six partials. Clean answers exist only at the edges. The 12 questions below 0.31 were all unanswerable, four of them off-topic. The five above 0.59 were all answerable.

Here's what each threshold would do if you hand off whenever confidence is below it:

| Threshold | Answerable, handed off (of 24) | Unanswerable, let through (of 35) | 
|---|---|---|
| 0.30 | 0 | 25 | 
| 0.35 | 1 | 17 | 
| 0.40 | 2 | 14 | 
| 0.45 | 7 | 9 | 
| 0.50 | 12 | 5 | 

There's no good row. At 0.40, 14 questions the retrieved text doesn't answer still go through, and 4 of the 6 partial ones are handed off. At 0.50, half of the answerable questions go to a human. The best single cut on this data, found after the fact, is about 0.43, and it still gets 11 of 59 wrong. I found it by looking at labels you won't have in production, and it doesn't travel: at 0.43, more than half of those 101 production answers would have gone to a human.

Two questions show the problem better than the tables.

**The highest-scoring "no":** "How many employees does Asktopus have?" scored 0.590. The site doesn't say. The top chunk was the features text about inviting colleagues as editors or viewers, which is about the customer's team, not ours. That score beats 19 of the 24 answerable questions.

**The lowest-scoring "yes":** "How long does it take to set up?" scored 0.309. The best match was a homepage chunk that ends with the FAQ question "How fast can I get started?". The answer to it falls in the next chunk, which wasn't retrieved. The answer came from the installation guide in second place: "up and running on your website in under 5 minutes". 22 of the 30 on-topic unanswerable questions scored higher.

An embedding captures what a text is about. Cosine similarity tells you the question and a chunk are about the same thing, not that the chunk contains the fact the question asks for. "How many employees does Asktopus have?" is about Asktopus and teams, and so is the features page. That's all the score can see.

The clearest pattern in the data: 30 of the unanswerable questions were on-topic. The 14 of them that use the product's own words ("Asktopus", "chatbot", "assistant", "widget", "AI") had a median score of 0.498, the same as the answerable questions. The 16 that don't had a median of 0.307. On a site about AI chatbots, any question about AI chatbots is close to everything.

What's inside the chunks makes it worse in both directions:

This is a small test. One small site, one crawl (partly out of date), 65 questions I wrote myself in English, labelled by my own judgement. Different content, languages or chunking would move the numbers. I also measured retrieval, not answers: I didn't test how reliably the model says "I don't know" when the chunks don't contain the answer. But nothing in the mechanism is specific to this site, and the production numbers point the same way.

In Asktopus, the score doesn't make any decision about the conversation.

In a first live check on our own site, the marker flagged all three questions the site doesn't cover (one of them in Danish) and neither a pricing question it does cover nor a plain "Hi there!". Five questions prove nothing statistically, but the signal comes from something that reads the text, which is the point.

We also stopped using confidence to pick the model. We used to send low-confidence questions to the bigger one, but a bigger model can't answer from content it wasn't given.

Asktopus, the chatbot this came out of, is at [asktopus.com](https://asktopus.com).
