My support chatbot scored 0.92. It was also lying to customers. A developer building a customer-support chatbot on Amazon Bedrock AgentCore for a Udacity × AWS Agent Engineer Nanodegree project found that a 0.92 LLM-as-a-judge evaluation score masked a critical failure: the bot fabricated a bug-report ticket ID (TIX-345678) that was never written to DynamoDB. The engineer traced the gap to the evaluation's inability to verify tool side effects, and now checks the database against chat transcripts instead of trusting the score alone. Twelve of my thirteen test prompts scored a perfect 1.00. Two independent evaluation runs agreed: 0.92 correctness . No message routed to the wrong place, both prompt-injection attempts refused. And yet, in one conversation, my chatbot told a customer: "I have filed a bug report with ticket ID TIX-345678." That ticket did not exist. Nothing had been written to the database. The customer would have walked away believing their bug was logged. This is the story of how I built that chatbot, why the score couldn't see the problem, and what I check now instead. I'm going through the Udacity × AWS Agent Engineer Nanodegree . The first project: a customer-support chatbot for a fictional online shop that handles three kinds of messages. One twist: Bedrock Agents Classic , which the course was originally designed around, closed to new customers on 30 July 2026. So the project runs on its successor, the Amazon Bedrock AgentCore managed harness . The other twist, and the actual point of the exercise: there's no classifier and no routing node. All of the routing, information-gathering, and grounding behaviour lives in a single system prompt. Customer ──► chat.py ──invoke harness──► AgentCore managed harness ◄──► Amazon Nova Pro │ system prompt.txt + FAQ temp 0, topK 1 │ tool call: bugreports create bug report ▼ AgentCore Gateway MCP, IAM auth ▼ Lambda create bug report ──PutItem──► DynamoDB The pieces: {{FAQ}} placeholder when the harness is created.