{"slug": "how-i-fixed-llm-counting-hallucinations-using-hindsight-facts", "title": "How I fixed LLM counting hallucinations using Hindsight facts", "summary": "A developer building a multi-channel customer support memory agent fixed LLM counting hallucinations by moving the count out of the prompt and into Python. The agent, built on FastAPI with Hindsight as the memory layer and Groq-hosted qwen/qwen3-32b, now has the model classify each interaction once at write time into structured facts (issue_id, channel, resolved), then recalls those facts and counts unresolved repeats deterministically to decide escalation. The developer reports the model previously undercounted rephrased complaints and overcounted resolved side questions, and that the split leaves the model only to write the explanation around a number it no longer produces.", "body_md": "The first time my support agent told me a customer had contacted us \"twice\" when the record showed four separate contacts, I assumed I had a retrieval bug. I didn't. The memory was fine. I had asked a language model to count, and it did what language models do when you ask them to count: it produced a confident, plausible number.\n\nThis is the story of how I stopped asking the model to count, and why Hindsight turned out to be the right place to put the thing that does.\n\nWhat the system does\n\nThe project is a customer support memory agent. A customer can reach us on chat, email, or phone, and nobody should have to explain their problem again when they switch channels. Everything is keyed on the customer's email address, which is the one identifier all three channels reliably share.\n\nThe backend is a small FastAPI service:\n\nhindsight_client.py wraps the memory layer, Hindsight, for storing and fetching interactions.\n\nllm_client.py wraps Groq, where I run qwen/qwen3-32b.\n\nmemory_schema.py holds the Pydantic models for an interaction.\n\napp.py exposes the endpoints, mainly POST /summarise/{email} and POST /escalate/{email}.\n\nThe summarise endpoint produces a short brief of a customer's whole history so the next agent, human or otherwise, starts informed. The escalate endpoint answers a narrower question: has this customer contacted us three or more times about the same unresolved issue? If yes, stop trying another generic fix and hand them to someone senior.\n\nThat second endpoint is where things went wrong.\n\nThe bug: asking a model to count\n\nMy first version of escalation was the obvious one. Recall the customer's memories, paste them into the prompt, and ask:\n\nBased on the history below, how many times has this customer contacted us about the same unresolved issue? If three or more, recommend escalation.\n\nIt looked fine on easy cases. It got shakier on the case I built specifically to stress it: a customer with a repeat billing problem spread across chat, email, and phone, with the issue described differently each time. The model would sometimes undercount because it treated a rephrased complaint as a new topic, and sometimes overcount because it counted a resolved side question as a repeat.\n\nThe pattern was consistent enough to stop being interesting. Counting is a deterministic operation. Deciding whether two messages are about the same issue is a judgement call. I had bundled both into one prompt and handed the whole thing to a component that is good at the second and unreliable at the first.\n\nI also couldn't debug it. When the model said \"two\", I had no intermediate value to inspect. The number was born inside the generation and vanished with it.\n\nThe fix: store facts, compute the number\n\nThe change was to split the work along the line between judgement and arithmetic.\n\nWhen an interaction happens, the model classifies it once and I store the result as structured facts in Hindsight: which issue it belongs to, its channel, and whether it was resolved.\n\nWhen I need a count, I recall those facts and count them in Python.\n\nThe model only receives the number, and only writes the explanation around it.\n\nThe shape of an interaction became something I could count over:\n\nfrom pydantic import BaseModel\n\nfrom typing import Literal\n\nclass Interaction(BaseModel):\n\n    email: str\n\n    channel: Literal[\"chat\", \"email\", \"phone\"]\n\n    issue_id: str          # normalised issue label, assigned once at write time\n\n    summary: str\n\n    resolved: bool\n\nThe important field is issue_id. It is the one place where the \"is this the same issue?\" judgement lives, and it is decided at write time, once, and then stored. Everything downstream reads it as a fact.\n\nRecall then returns those stored facts for a customer, and the counting is boring on purpose:\n\nfrom collections import Counter\n\ndef open_issue_counts(interactions: list[Interaction]) -> Counter:\n\n    return Counter(i.issue_id for i in interactions if not i.resolved)\n\ndef needs_escalation(interactions: list[Interaction], threshold: int = 3) -> bool:\n\n    counts = open_issue_counts(interactions)\n\n    return any(n >= threshold for n in counts.values())\n\nThere is nothing clever here, and that is the point. A Counter does not hallucinate. It also gives me something I never had before: an intermediate value I can log, assert on, and put in a unit test.\n\nWhere Hindsight earns its place\n\nI could have kept these facts in a plain Postgres table, and for the counting alone that would have worked. What made me keep Hindsight in the middle is that the summarise endpoint and the escalate endpoint need different things from the same history.\n\nEscalation needs exact, structured records. Summarisation needs the relevant parts of a messy history, ranked by what matters right now. Retrieval that can serve both from one store meant I didn't have to maintain a relational table for the first job and a vector index for the second and keep them in sync. If you want the background on why this distinction matters, Vectorize's explanation of agent memory is the clearest write-up I've found on memory that goes beyond replaying chat history.\n\nThe wrapper stays thin so the rest of the code doesn't know or care how retrieval works:\n\ndef recall_interactions(email: str) -> list[Interaction]:\n\n    raw = hindsight.recall(subject=email)      # everything remembered for this customer\n\n    return [Interaction(**item) for item in raw]\n\ndef retain_interaction(interaction: Interaction) -> None:\n\n    hindsight.retain(subject=interaction.email, content=interaction.model_dump())\n\nKeying on email is the decision that makes the cross-channel behaviour work at all. A chat session ID would have fragmented the history by channel, and the whole product depends on the three channels landing in one place.\n\nWhat the endpoints do now\n\nThe escalation endpoint no longer asks the model anything it can't be trusted with:\n\n@app.post(\"/escalate/{email}\")\n\ndef escalate(email: str):\n\n    interactions = recall_interactions(email)\n\n    counts = open_issue_counts(interactions)\n\n    escalate_now = needs_escalation(interactions)\n\n```\nexplanation = llm.explain_escalation(\n    counts=dict(counts),      # the model is handed the numbers, not asked for them\n    escalate=escalate_now,\n    history=interactions,\n)\nreturn {\"escalate\": escalate_now, \"counts\": dict(counts), \"explanation\": explanation}\n```\n\nThe response carries counts alongside the explanation, so anyone reading the output can check the model's prose against the number it was given.\n\nTake the seed customer I use for the escalation case: four interactions across chat, email, and phone, all about the same unresolved billing problem. The count for that issue_id is four, the threshold check fires, and the explanation reads as a summary of a fact rather than a guess at one. The customer with a bug that was reported and resolved with a workaround produces no open issues and no escalation, which is the behaviour I want from a false-positive test.\n\nThe model still does real work. It writes the brief, and it explains why a case is being escalated in language a support lead can read in ten seconds. It just no longer gets to decide the facts the explanation rests on.\n\nLessons learned\n\nDon't ask a model to do arithmetic you can do in code. If the answer is a count, a sum, a date difference, or a set membership check, compute it. Give the model the result and ask it to communicate it.\n\nPut judgement at write time, not read time. Deciding \"same issue or not\" once, when an interaction is stored, is cheaper and more consistent than re-deciding it every time someone asks a question. It also means the decision is stored where you can inspect and correct it.\n\nA number the model never produced is a number you can test. The moment counts came from a Counter, I could write assertions against them. Before that, my only test was reading the output and deciding whether it seemed right.\n\nShow the evidence next to the explanation. Returning counts in the API response costs nothing and lets a human catch a bad explanation immediately.\n\nBe honest about the part that is still fuzzy. The issue_id assignment is still a model call, and it can still be wrong. I moved the uncertainty into one place I can watch, rather than eliminating it. A rephrased complaint that gets a new issue_id will still evade the threshold, and that is the next thing I want to tighten.\n\nThe general shape of this fix is worth borrowing even if you never touch a support tool: find the operations in your agent that are secretly arithmetic, and take them away from the model. The Hindsight docs are a good place to start if you want a memory layer that can hold the facts while you do it.", "url": "https://wpnews.pro/news/how-i-fixed-llm-counting-hallucinations-using-hindsight-facts", "canonical_source": "https://dev.to/sri_varsha_527/how-i-fixed-llm-counting-hallucinations-using-hindsight-facts-25n1", "published_at": "2026-09-29 03:38:36+00:00", "updated_at": "2026-09-29 03:46:44.123666+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "mlops"], "entities": ["Hindsight", "FastAPI", "Groq", "qwen/qwen3-32b", "Pydantic"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-i-fixed-llm-counting-hallucinations-using-hindsight-facts", "markdown": "https://wpnews.pro/news/how-i-fixed-llm-counting-hallucinations-using-hindsight-facts.md", "text": "https://wpnews.pro/news/how-i-fixed-llm-counting-hallucinations-using-hindsight-facts.txt", "jsonld": "https://wpnews.pro/news/how-i-fixed-llm-counting-hallucinations-using-hindsight-facts.jsonld"}}