{"slug": "retrieval-confidence-can-t-tell-your-rag-chatbot-when-the-answer-is-missing", "title": "Retrieval confidence can't tell your RAG chatbot when the answer is missing", "summary": "A developer at Asktopus tested whether retrieval confidence scores can reliably tell a RAG chatbot when its knowledge base lacks an answer, running 65 questions against the company's own 44-page knowledge base split into 219 chunks. The experiment found that answerable and unanswerable questions overlap heavily in cosine similarity (0.31 to 0.59), so no single threshold works: at the documented 0.4 cutoff, 14 of 35 unanswerable questions still pass through, while at 0.5 half of the answerable ones are wrongly handed to a human. The best post-hoc cut of about 0.43 still misclassifies 11 of 59 questions and would have sent more than half of 101 real production answers to a human.", "body_md": "*I ran 65 questions against our own knowledge base and checked which ones the retrieved text actually answered. The similarity scores overlapped too much for any threshold to work.*\n\nOur own handoff docs used to describe it: if retrieval confidence falls below a threshold, hand the visitor to a human instead of generating a weak answer. The default was 0.4. It's the design most people building a RAG chatbot reach for, and I no longer think it works. This post is about why: first a hunch from production, then a small experiment.\n\nIn a typical RAG pipeline, including ours at Asktopus, the visitor's question is embedded (we use OpenAI's text-embedding-3-small). The vector database (Qdrant) returns the five closest chunks by cosine similarity, and those chunks go into the prompt. The best of those five scores is what usually gets called \"retrieval confidence\".\n\nIt's free, already computed and between 0 and 1, so it looks like the system telling you how sure it is.\n\n```\nif confidence < THRESHOLD:\n    hand_off_to_human()      # or reply \"I don't know\"\nelse:\n    answer_from_context()\n```\n\nIt promises two things at once: no made-up answers when the knowledge is thin, and a human for the questions the bot can't handle. All you have to do is pick the number.\n\nIn production, the scores of 101 real answers on one site ranged from 0.11 to 0.67, with a median of 0.36. At the 0.4 our docs recommended, more than half of them would have gone to a human. I couldn't tell from the numbers which of those answers were good, and 101 answers on one site is far too few to tune a threshold on. So I set up a test where I could check.\n\nThe knowledge base is asktopus.com itself: 44 crawled pages in English and Danish, split into 219 chunks of up to 1,500 characters. The pipeline is exactly production's: same embedding model, same Qdrant search, top 5, confidence = the best score.\n\nI wrote 65 questions: 40 I expected the site to answer, 20 I expected it not to (a Zapier integration, ISO 27001, an uptime SLA), and 5 plainly off-topic (the capital of France, tomorrow's weather). Then I set my expectations aside and labelled each question by reading the five chunks the model would receive. Does this text contain the answer? **Yes**, **partial** (it answers a narrower version: conversations can be exported, but the format isn't stated), or **no**.\n\nMy expectations were off. Of the 40 questions I expected the site to answer, only 24 got retrieved text that answered them; 6 were partial and 10 got nothing useful. Three of those ten were retrieval misses: the answer was in the knowledge base, just not in the top five. All 25 questions I expected it not to answer got a no.\n\n| Label | Questions | Min | Median | Max | \n|---|---|---|---|---|\n| Yes | 24 | 0.309 | 0.499 | 0.708 | \n| Partial | 6 | 0.331 | 0.379 | 0.582 | \n| No | 35 | 0.102 | 0.345 | 0.590 | \n\nThe score isn't noise: answerable questions score higher on average. But the two groups overlap from 0.31 to 0.59, and that band holds 19 of the 24 answerable questions, 23 of the 35 unanswerable ones and all six partials. Clean answers exist only at the edges. The 12 questions below 0.31 were all unanswerable, four of them off-topic. The five above 0.59 were all answerable.\n\nHere's what each threshold would do if you hand off whenever confidence is below it:\n\n| Threshold | Answerable, handed off (of 24) | Unanswerable, let through (of 35) | \n|---|---|---|\n| 0.30 | 0 | 25 | \n| 0.35 | 1 | 17 | \n| 0.40 | 2 | 14 | \n| 0.45 | 7 | 9 | \n| 0.50 | 12 | 5 | \n\nThere's no good row. At 0.40, 14 questions the retrieved text doesn't answer still go through, and 4 of the 6 partial ones are handed off. At 0.50, half of the answerable questions go to a human. The best single cut on this data, found after the fact, is about 0.43, and it still gets 11 of 59 wrong. I found it by looking at labels you won't have in production, and it doesn't travel: at 0.43, more than half of those 101 production answers would have gone to a human.\n\nTwo questions show the problem better than the tables.\n\n**The highest-scoring \"no\":** \"How many employees does Asktopus have?\" scored 0.590. The site doesn't say. The top chunk was the features text about inviting colleagues as editors or viewers, which is about the customer's team, not ours. That score beats 19 of the 24 answerable questions.\n\n**The lowest-scoring \"yes\":** \"How long does it take to set up?\" scored 0.309. The best match was a homepage chunk that ends with the FAQ question \"How fast can I get started?\". The answer to it falls in the next chunk, which wasn't retrieved. The answer came from the installation guide in second place: \"up and running on your website in under 5 minutes\". 22 of the 30 on-topic unanswerable questions scored higher.\n\nAn embedding captures what a text is about. Cosine similarity tells you the question and a chunk are about the same thing, not that the chunk contains the fact the question asks for. \"How many employees does Asktopus have?\" is about Asktopus and teams, and so is the features page. That's all the score can see.\n\nThe clearest pattern in the data: 30 of the unanswerable questions were on-topic. The 14 of them that use the product's own words (\"Asktopus\", \"chatbot\", \"assistant\", \"widget\", \"AI\") had a median score of 0.498, the same as the answerable questions. The 16 that don't had a median of 0.307. On a site about AI chatbots, any question about AI chatbots is close to everything.\n\nWhat's inside the chunks makes it worse in both directions:\n\nThis is a small test. One small site, one crawl (partly out of date), 65 questions I wrote myself in English, labelled by my own judgement. Different content, languages or chunking would move the numbers. I also measured retrieval, not answers: I didn't test how reliably the model says \"I don't know\" when the chunks don't contain the answer. But nothing in the mechanism is specific to this site, and the production numbers point the same way.\n\nIn Asktopus, the score doesn't make any decision about the conversation.\n\nIn a first live check on our own site, the marker flagged all three questions the site doesn't cover (one of them in Danish) and neither a pricing question it does cover nor a plain \"Hi there!\". Five questions prove nothing statistically, but the signal comes from something that reads the text, which is the point.\n\nWe also stopped using confidence to pick the model. We used to send low-confidence questions to the bigger one, but a bigger model can't answer from content it wasn't given.\n\nAsktopus, the chatbot this came out of, is at [asktopus.com](https://asktopus.com).", "url": "https://wpnews.pro/news/retrieval-confidence-can-t-tell-your-rag-chatbot-when-the-answer-is-missing", "canonical_source": "https://dev.to/klausbyskov/retrieval-confidence-cant-tell-your-rag-chatbot-when-the-answer-is-missing-2ml0", "published_at": "2026-10-08 07:38:21+00:00", "updated_at": "2026-10-08 07:48:32.360979+00:00", "lang": "en", "topics": ["ai-search", "large-language-models", "ai-products", "natural-language-processing"], "entities": ["Asktopus", "OpenAI", "text-embedding-3-small", "Qdrant", "Zapier"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/retrieval-confidence-can-t-tell-your-rag-chatbot-when-the-answer-is-missing", "markdown": "https://wpnews.pro/news/retrieval-confidence-can-t-tell-your-rag-chatbot-when-the-answer-is-missing.md", "text": "https://wpnews.pro/news/retrieval-confidence-can-t-tell-your-rag-chatbot-when-the-answer-is-missing.txt", "jsonld": "https://wpnews.pro/news/retrieval-confidence-can-t-tell-your-rag-chatbot-when-the-answer-is-missing.jsonld"}}