Node.js Evidence Lifecycles: 5 RAG Hallucination Fixes When Docs Chatbot Answers Go Wrong A developer outlines a versioned evidence pipeline for ask-your-docs RAG systems, arguing that hallucination in code-review chatbots stems from stale, incomplete, or conflicting retrieved evidence rather than weak embeddings or small context windows. The approach pins repository and course revisions, retrieves code and policy evidence through separate lanes, rejects weak evidence, validates findings against a typed schema, and replays a fixed evaluation set before release. The developer recommends owning the evidence contract in Node.js and treating the model provider as an adapter. Short answer: RAG hallucination persists when an ask-your-docs chatbot receives stale, incomplete, or conflicting evidence, so good embeddings and a larger context window still produce wrong answers. For an edtech assistant that reviews code changes, the practical fix is a versioned evidence pipeline: pin the course and repository revision, retrieve requirements and changed code separately, reject weak evidence, validate a small finding schema, and replay a fixed evaluation set before release. | Choice | Wrong-answer risk | Provider portability | Operating burden | |---|---|---|---| | Send the nearest chunks directly to generation | High when revisions or document types conflict | Low if prompts depend on one provider's response shape | Low | | Add reranking but no evidence gate | Lower ranking risk; unsupported output can still escape | Medium | Medium | | Use a typed evidence contract plus abstention | Explicitly controlled at the application boundary | High | Medium | Recommendation: own the evidence contract in Node.js and make the model provider an adapter. This costs more engineering time than a single prompt, but it protects the scarce resource in a one-person SaaS: hours that should go into the weekly product release, not manual review of plausible findings. Embeddings answer a narrow question: which indexed passages appear related to the query? A code-review assistant needs a harder answer: which passages are authoritative for this exact change, repository revision, assignment, language, and policy version? Those are different jobs. Imagine a learner changes gradeSubmission in revision 8f21c4a . The index contains the current rubric, last term's rubric, an instructor exception, generated API documentation, and a discussion that quotes obsolete behavior. A semantically close chunk can be irrelevant because its authority or effective revision is wrong. Increasing the context window may make matters worse: both rules now fit, so generation receives more contradiction rather than more truth. My rule is blunt: freshness is a hard filter, not a similarity hint. Every indexed unit should carry fields such as repositoryId , revision , documentKind , effectiveFrom , and supersedesId , plus stable source coordinates. A finding that cites “somewhere in the handbook” isn't reviewable; one tied to a file, symbol, rubric item, and immutable content digest is. I won't trade that traceability for a marginally simpler ingestion job because the later debugging cost lands on the same person trying to ship the next feature. Version first. This changes the debugging question. Stop asking only, “Was the right chunk in the top five?” Ask, “Was every eligible chunk valid for the reviewed revision, and did an older rule survive ingestion?” The second question catches a class of failures that embedding swaps and larger windows cannot. A diff and a rubric play different roles. The diff describes what changed. The rubric defines what should be true. Mixing both into one vector query produces a convenient bag of text, but it erases that distinction before the model has to reason about it. Use two retrieval lanes. The first resolves code evidence from the changed files, surrounding symbols, tests, and pinned base revision. The second resolves policy evidence from the active assignment rubric and applicable instructor rules. Join them only after each lane passes metadata filters. This is intentionally boring infrastructure. Boring ships. The join should create explicit claim candidates such as “the new branch can return an ungraded state” paired with code coordinates and the rubric clause that makes the state invalid. If either side is missing, there is no finding yet. There is only a search lead. Chunk boundaries matter here, but fixed token counts are a blunt default. Keep a function with its signature, keep a rubric requirement with its exceptions, and preserve parent headings as metadata. Overlap can help a boundary case, though excessive overlap duplicates evidence and can make several retrieved chunks look like independent support. Deduplicate by source identity and content digest before generation. For a weekly release cadence, I would outsource parsing where a maintained parser already exists and keep the domain rules in application code. The differentiator is not splitting Markdown. It is knowing that “late submission exception” overrides the general deadline only for a named assignment cohort. Similarity scores are not calibrated truth probabilities. Do not print a universal cutoff into a blog post and pretend it transfers across embedding models, corpora, and distance functions. Choose thresholds from labeled queries in your own corpus, then version those thresholds alongside the retriever. The gate still needs deterministic rules. Here is a compact contract for an initial implementation: type Evidence = { sourceId: string; digest: string; revision: string; kind: "diff" | "code" | "rubric" | "policy" | "test"; locator: string; score: number; text: string; }; type ReviewRequest = { repositoryId: string; revision: string; rubricRevision: string; }; type EvidenceBundle = { code: Evidence ; rules: Evidence ; }; function admitEvidence request: ReviewRequest, candidates: Evidence , minimumScore: number, : EvidenceBundle | null { const eligible = candidates.filter item = item.score = minimumScore && item.kind === "rubric" || item.kind === "policy" ? item.revision === request.rubricRevision : item.revision === request.revision , ; const unique = ...new Map eligible.map item = ${item.sourceId}:${item.digest} , item , .values ; const bundle = { code: unique.filter item = item.kind === "diff" || item.kind === "code" || item.kind === "test" , rules: unique.filter item = item.kind === "rubric" || item.kind === "policy" , }; return bundle.code.length 0 && bundle.rules.length 0 ? bundle : null; } The minimumScore is configuration derived from evaluation, not a magic number hidden in the function. The revision equality checks are more important than another decimal place of similarity. If admitEvidence returns null , the correct response is a typed insufficient evidence result. Do not ask the generator to improvise. That refusal path is product behavior, so design it deliberately. It can request the missing rubric, wait for indexing, or route the change to human review. OWASP's guidance for LLM applications treats excessive agency and overreliance as risk areas; keeping authorization and consequential actions outside generated prose follows the same defensive boundary. Free-form review text hides retrieval defects. A typed result exposes them. Require each finding to name the code location, the violated rule, supporting source identifiers, and a confidence category defined by your application. Then validate the object before it reaches a pull request or learner dashboard. JSON Schema is useful here because validation stays outside the model and does not depend on a provider-specific SDK. type Finding = { summary: string; severity: "info" | "warning" | "error"; codeLocator: string; ruleLocator: string; evidenceIds: string ; }; type ReviewResult = | { status: "complete"; findings: Finding } | { status: "insufficient evidence"; missing: string }; function evidenceIsClosed result: ReviewResult, admittedIds: ReadonlySet