{"slug": "my-assistant-answered-a-question-without-supporting-evidence", "title": "My assistant answered a question without supporting evidence", "summary": "A retrieval-augmented assistant built on GPT-5.6 Luna answered a Quadient Exstream PDF/A-3 configuration question three out of four times without any supporting document in its knowledge base, according to an August comparison of three models. In the same test of 12 unanswerable questions run four times per model, a checker marked 12 of 12 runs as refusals for GPT-5.6 Terra, 10 of 12 for GPT-4o, and seven of 12 for GPT-5.6 Luna. The author noted the checker matched refusal phrases and counted empty answers as refusals, so the scores do not prove every refusal was a clean stop, and the assistant should state when no supporting evidence exists and stop so the user can decide whether to escalate to a human.", "body_md": "I asked my Exstream assistant which Quadient setting controlled PDF/A-3 attachments. There was no supporting document in its knowledge base. It named a setting anyway.\n\nThis was one of the questions I used to compare three models in August. I wanted the assistant to answer from the documents I had given it, and stop when those documents could not support an answer.\n\nThe first test asked the same 12 Exstream support questions in both English and Norwegian. All three models found the expected documents. Finding them did not establish that every sentence in the resulting answers was correct, but it was a useful check of retrieval across languages.\n\nThe second test asked about things the documents did not cover. Alongside the Quadient question, I asked about Exstream certificate rotation and the cargo capacity of a Boeing 747-8F. Each question ran four times per model.\n\nThe checker marked 12 of 12 runs as refusals for GPT-5.6 Terra, 10 of 12 for GPT-4o, and seven of 12 for GPT-5.6 Luna. My notes record Luna answering the Quadient question three times out of four without a supporting document.\n\nI chose Terra at the time. The checker matched refusal phrases and counted empty answers as refusals. A response could mention escalation and still contain unsupported advice, so the scores did not prove that every refusal was a clean stop.\n\nNor did the Quadient example prove that the setting was wrong. That was a separate question. Even if the answer happened to be correct, it broke the requirement I had set for this assistant.\n\nThe test should include questions the documents cannot answer. Repeat them and read every response, including those the checker marked as refusals. A refusal phrase does not prove that the model stopped, because it can still give unsupported instructions after saying it lacks information.\n\nWhen there is no supporting evidence, my assistant should say so and stop. The user can then decide whether to escalate to a human.", "url": "https://wpnews.pro/news/my-assistant-answered-a-question-without-supporting-evidence", "canonical_source": "https://hajek.no/posts/2026/my-assistant-answered-without-evidence/", "published_at": "2026-09-22 05:53:32+00:00", "updated_at": "2026-09-22 06:24:53.541918+00:00", "lang": "en", "topics": ["ai-products", "large-language-models", "ai-tools"], "entities": ["GPT-5.6 Luna", "GPT-5.6 Terra", "GPT-4o", "Quadient", "Exstream", "Boeing 747-8F"], "alternates": {"html": "https://wpnews.pro/news/my-assistant-answered-a-question-without-supporting-evidence", "markdown": "https://wpnews.pro/news/my-assistant-answered-a-question-without-supporting-evidence.md", "text": "https://wpnews.pro/news/my-assistant-answered-a-question-without-supporting-evidence.txt", "jsonld": "https://wpnews.pro/news/my-assistant-answered-a-question-without-supporting-evidence.jsonld"}}