// double-click race test v5 // QA regression test marker
Deterministic post-response grounding checks + cross-family adversarial review for LLM chat agents — zero deps, bring your own LLM and your own database.
Quick start · Why · Docs · Compare · Roadmap
This is an anonymized real production case — a chat agent claimed a
file existed that didn't. Here's ora-grounding catching it, deterministically, with zero LLM calls in the check itself:
>>> from ora_grounding.grounding import extract_claims, classify_claims
>>>
>>> reply = "Fixed the retry logic in payments_client.py — added dedup via redis_lock.py"
>>>
>>> canonical = {
... "paths": {"src/payments_client.py"}, # redis_lock.py does NOT exist
... "basenames": {"payments_client.py"},
... "defs": set(),
... }
>>>
>>> classify_claims(extract_claims(reply), canonical=canonical)
{'fabricated': ['redis_lock.py'], 'unverified': []}
One real file. One invented file. Caught instantly. That's the whole pitch — everything below is detail.
LLMs hallucinate confidently. Two failure modes hurt users the most:
| Failure mode | What it looks like |
|---|---|
| Made-up specifics | "Fix at services/auth.py:42 " — the file doesn't exist. |
| Overconfident synthesis | The model stitches together plausible claims nothing in its context supports. |
Prompting alone doesn't fix this. Sibling-model review doesn't fix it either — GPT reviewing GPT shares blind spots. ora-grounding adds two deterministic defences that sit outside the model:
- 🧮 Cheap grounding check — regex + set-membership,no LLM in the hot path . Catches file/symbol/line-number/command claims the retrieval context never supported.
- 🥊 Adversarial review — adifferent-family reviewer LLM hostile-reads the draft, with a hard deterministic guard against the reviewer itself hallucinating flags.
Extracted from a production AI-CTO assistant serving real users. Battle-tested against actual regressions — including the one above.
pip install ora-grounding
Zero dependencies. Python 3.10+.
from ora_grounding.grounding import extract_claims, classify_claims
reply = "Fixed auth in backend/routers/auth.py line 42"
canonical = {
"paths": {"backend/routers/auth.py"},
"basenames": {"auth.py"},
"defs": {"verify_token", "login"},
}
claims = extract_claims(reply)
result = classify_claims(claims, canonical=canonical)
if result["fabricated"]:
print(f"⚠️ Fabricated: {result['fabricated']}")
python
from ora_grounding.review import adversarial_review
review_result = adversarial_review(
draft_reply=reply,
retrieval_context=your_rag_chunks,
reviewer_llm=your_llm_client, # Different family from the drafter
)
if review_result["flags"]:
print(f"🚩 Review flags: {review_result['flags']}")
| Approach | Speed | Catches fabricated paths | Catches overconfident synthesis | Cross-family |
|---|---|---|---|---|
| Prompting alone | Fast | ❌ | ❌ | N/A |
| Sibling-model review | Slow | ❌ | ||
| ora-grounding | Fast (grounding) + Slow (review) | ✅ | ✅ | ✅ |
-
Prompting alone — "Be accurate. Don't hallucinate." — doesn't work. The model doesn'tknow it's hallucinating.
-
Sibling-model review — GPT-4 reviewing GPT-4 shares blind spots. Same training data, same failure modes.
-
ora-grounding — Deterministic check (fast) + adversarial review (slow, opt-in) with a different-family reviewer.
-
Deterministic grounding check
-
Adversarial review with cross-family LLM
-
Structured output validation (Pydantic models)
-
Multi-turn conversation grounding
-
Benchmark suite (public dataset)
MIT — see LICENSE.
Extracted from AUREM — an AI-CTO assistant that reads your GitHub repo and ships code. Built by the AUREM team.
Questions? Open an issue or reach out at support@aurem.com.