# ClauseHound: a pet-insurance decision engine that refuses to guess

> Source: <https://dev.to/ling_zhou_78cf1e40f3e605d/clausehound-a-pet-insurance-decision-engine-that-refuses-to-guess-14f>
> Published: 2026-10-01 19:35:27+00:00

*This is a submission for the [Sanity Challenge](https://dev.to/challenges/sanity-2026-09-16), Path One: Ship an Agent That Queries Real Content.*

🔗 **Demo:** [https://clausehound-1rmpnbeb5-lingzhoudesign-gmailcoms-projects.vercel.app](https://clausehound-1rmpnbeb5-lingzhoudesign-gmailcoms-projects.vercel.app)

🗄️ **Sanity project ID:** `yijqzehr` (dataset `production`)

🏷️ **Tag:** `#sanitychallenge`

Your Bernese Mountain Dog needs hip surgery at age three. You file the claim, confident — and get denied. The 12-month orthopedic waiting period was in Section 4 all along. You just never found it *before* you bought the policy.

That's the moment ClauseHound is built for — except *before* it happens. **ClauseHound is a pet-insurance decision engine for people shopping for a policy.** Not a chatbot that answers trivia about pet insurance: it takes the four decisions shoppers actually agonize over and answers each one from the insurers' own policy text, with the exact clause cited:

One grounding contract runs through all four:

`unresolved`. The agent never picks a winner.
No login needed. The landing page frames the four decisions; each job card (and the hero) opens the chat modal with its question already prefilled — you press Send, it never auto-sends.

Try these in order:

*The landing page frames the four decisions, not a chat box.*

*Each job card opens the modal with its question prefilled. Nothing auto-sends.*

*The answered state: "There's no single winner — it depends on when the emergency hits." Verdict first, then the month-by-month math, then the cited clauses.*

Next.js + TypeScript app. The interesting parts:

`app/lib/mcp.ts` — Sanity Context MCP client (JSON-RPC over HTTPS). The app never queries Sanity directly; every fact comes through the Context endpoint's tools (`initial_context`, `groq_query`, …).` app/lib/cost.ts` — deterministic money math in integer cents: premiums, deductibles, reimbursement rates, annual limits, multi-year scenarios. The LLM explains the numbers; it never computes them. (Tested to the cent — `cost.test.ts`.)` app/lib/insureVsSave.ts`, `switching.ts`, `denial.ts`, `renewals.ts`, `appeals.ts` — the four decision jobs as typed modules, each with its own test suite.`app/components/ChatModal.tsx` + `ChatApp.tsx` — the modal chat: prefilled questions, streaming answers with loading stages, citation cards, contradiction panels, cost cards.
The repo is private (`lymcho/clausehound`, now merged to `main`); the test suite (1,653 tests) and the eval fixtures are part of the submission evidence below.

Everything ClauseHound knows lives in Sanity as structured content — not blobs of text, but typed documents with relationships, modeled from the insurers' published US sample policies:

`production` dataset, verified live.
The schema is the trust strategy. Every clause document carries a required `sectionRef` (e.g. `Sec. 3.4`) — the citation anchor — and a required `policy` reference back to its `policyDocument`, which references its `insurer`:

```
insurer → policyDocument → coverageClause / exclusionClause / waitingPeriod / claimStep
                          (each with required sectionRef)
```

The agent reads through a **Sanity Context MCP endpoint** backed by an Agent Context over that dataset. A GROQ content filter at the context level scopes every query to the four answerable types and only documents attached to a real policy:

```
_type in ["coverageClause", "exclusionClause", "waitingPeriod", "claimStep"]
&& defined(*[_type == "policyDocument" && _id == ^.policy._ref][0])
```

That filter is the guardrail: the agent *cannot* wander into irrelevant content no matter what the user asks. Coverage, exclusions, and waiting periods are distinct document types and stay distinct in answers — which is exactly why the conflict view works: a `coverageClause` and an `exclusionClause` matching the same question surface as two typed records with their `sectionRef` s, not as two paragraphs to blend.

The app runs an agentic loop over the MCP tools and narrates the pipeline at real phase boundaries — "Pulling the relevant policy text…", "Analyzing your question against the fine print…", "Merging the insurer-by-insurer findings…", "Verifying every claim against the policy text…". The grounding isn't just enforced, it's *visible* while you wait.

This is the bet the challenge asks for: **the agent only works because the content was structured.** Keyword search over the same policies could find the words "hip dysplasia"; it could not keep five carriers' answers from bleeding into each other, keep a coverage clause distinct from the exclusion that overrides it, or compute a 10-year cost from typed premium/deductible/reimbursement fields. The structure *is* the product.

The release is covered by a 19-fixture held-out eval (`eval/fixtures.json`, fixtures frozen — never edited to make the app pass). Each fixture posts a real user question to the deployed `/api/chat` and checks the response structurally: HTTP 200, citations resolve to real clause IDs, no dangling markers, carrier scope correct, refusals refuse, contradictions surface as `unresolved`. The two cost fixtures go further: they paste real quotes and re-derive every dollar figure with an independent port of the integer-cent math — exact match required, no LLM judging.

**16/19 passed** on the production build, 2026-10-01 (merged to `main` as `9957cb1`; the merge tree is byte-identical to the evaluated feature commit `6f4e452b`, so the results stand for production). Live-answer latencies ran 18–95s (streamed); the deterministic cost fixtures verify in under a second.

Honest failures, since they matter more than the score:

Known limitations the eval doesn't hide:

`yijqzehr`, dataset `insurer`, `policyDocument`, `coverageClause`, `exclusionClause`, `waitingPeriod`, `claimStep`
`SANITY_AGENT_CONTEXT_URL`), backed by an Agent Context over the production dataset; the GROQ content filter above scopes all retrieval to answerable, policy-attached documents.
No login is needed to try ClauseHound. Live answers are streamed; the deterministic cost comparisons run without the model.

*Demo only — research aid, not professional insurance advice. Verify with your insurer before buying.*
