Deterministic grounding checks for LLM agents, no LLM in the hot path The open-source Python package ora-grounding, released under the MIT license and installable via `pip install ora-grounding`, adds deterministic post-response grounding checks plus cross-family adversarial review for LLM chat agents with zero dependencies and no LLM calls in the grounding check itself. The tool, extracted from a production AI-CTO assistant, uses regex and set-membership to flag fabricated file, symbol, line-number, and command claims — in its demo it caught an invented `redis_lock.py` while confirming the real `payments_client.py` — and pairs that with a reviewer LLM from a different model family to catch overconfident synthesis. The package requires Python 3.10+ and targets hallucination failure modes that prompting alone and sibling-model review do not fix. // double-click race test v5 // QA regression test marker Deterministic post-response grounding checks + cross-family adversarial review for LLM chat agents — zero deps, bring your own LLM and your own database. Quick start -30-second-demo · Why -why-this-exists · Docs -usage · Compare -vs-the-alternatives · Roadmap -roadmap This is an anonymized real production case — a chat agent claimed a file existed that didn't. Here's ora-grounding catching it, deterministically, with zero LLM calls in the check itself: python from ora grounding.grounding import extract claims, classify claims reply = "Fixed the retry logic in payments client.py — added dedup via redis lock.py" canonical = { ... "paths": {"src/payments client.py"}, redis lock.py does NOT exist ... "basenames": {"payments client.py"}, ... "defs": set , ... } classify claims extract claims reply , canonical=canonical {'fabricated': 'redis lock.py' , 'unverified': } One real file. One invented file. Caught instantly. That's the whole pitch — everything below is detail. LLMs hallucinate confidently . Two failure modes hurt users the most: | Failure mode | What it looks like | |---|---| | Made-up specifics | "Fix at services/auth.py:42 " — the file doesn't exist. | | Overconfident synthesis | The model stitches together plausible claims nothing in its context supports. | Prompting alone doesn't fix this. Sibling-model review doesn't fix it either — GPT reviewing GPT shares blind spots. ora-grounding adds two deterministic defences that sit outside the model: - 🧮 Cheap grounding check — regex + set-membership, no LLM in the hot path . Catches file/symbol/line-number/command claims the retrieval context never supported. - 🥊 Adversarial review — a different-family reviewer LLM hostile-reads the draft, with a hard deterministic guard against the reviewer itself hallucinating flags. Extracted from a production AI-CTO assistant serving real users. Battle-tested against actual regressions — including the one above. pip install ora-grounding Zero dependencies. Python 3.10+. python from ora grounding.grounding import extract claims, classify claims 1. Your agent generates a reply reply = "Fixed auth in backend/routers/auth.py line 42" 2. Build the canonical set from your retrieval context canonical = { "paths": {"backend/routers/auth.py"}, "basenames": {"auth.py"}, "defs": {"verify token", "login"}, } 3. Check claims = extract claims reply result = classify claims claims, canonical=canonical if result "fabricated" : print f"⚠️ Fabricated: {result 'fabricated' }" python from ora grounding.review import adversarial review After the grounding check passes, run a cross-family review review result = adversarial review draft reply=reply, retrieval context=your rag chunks, reviewer llm=your llm client, Different family from the drafter if review result "flags" : print f"🚩 Review flags: {review result 'flags' }" | Approach | Speed | Catches fabricated paths | Catches overconfident synthesis | Cross-family | |---|---|---|---|---| | Prompting alone | Fast | ❌ | ❌ | N/A | | Sibling-model review | Slow | | | ❌ | | ora-grounding | Fast grounding + Slow review | ✅ | ✅ | ✅ | - Prompting alone — "Be accurate. Don't hallucinate." — doesn't work. The model doesn't know it's hallucinating. - Sibling-model review — GPT-4 reviewing GPT-4 shares blind spots. Same training data, same failure modes. - ora-grounding — Deterministic check fast + adversarial review slow, opt-in with a different-family reviewer. - Deterministic grounding check - Adversarial review with cross-family LLM - Structured output validation Pydantic models - Multi-turn conversation grounding - Benchmark suite public dataset MIT — see LICENSE https://github.com/polarisbuiltinc-wq/ora-grounding/blob/main/LICENSE . Extracted from AUREM https://aurem.com — an AI-CTO assistant that reads your GitHub repo and ships code. Built by the AUREM team. Questions? Open an issue or reach out at support@aurem.com mailto:support@aurem.com .