ClauseHound: a pet-insurance decision engine that refuses to guess A developer built ClauseHound, a pet-insurance decision engine that answers four shopper questions β€” coverage, switching, denial, renewals and appeals β€” strictly from insurers' published policy text, citing the exact clause via a required sectionRef anchor. The Next.js and TypeScript app routes every fact through a Sanity Context MCP endpoint with a GROQ filter that scopes queries to coverageClause, exclusionClause, waitingPeriod and claimStep documents, and it returns 'unresolved' rather than picking a winner when clauses conflict. All money math is deterministic integer-cent code in app/lib/cost.ts, with the LLM only explaining numbers it never computes, backed by a 1,653-test suite. This is a submission for the Sanity Challenge https://dev.to/challenges/sanity-2026-09-16 , Path One: Ship an Agent That Queries Real Content. πŸ”— Demo: https://clausehound-1rmpnbeb5-lingzhoudesign-gmailcoms-projects.vercel.app https://clausehound-1rmpnbeb5-lingzhoudesign-gmailcoms-projects.vercel.app πŸ—„οΈ Sanity project ID: yijqzehr dataset production 🏷️ Tag: sanitychallenge Your Bernese Mountain Dog needs hip surgery at age three. You file the claim, confident β€” and get denied. The 12-month orthopedic waiting period was in Section 4 all along. You just never found it before you bought the policy. That's the moment ClauseHound is built for β€” except before it happens. ClauseHound is a pet-insurance decision engine for people shopping for a policy. Not a chatbot that answers trivia about pet insurance: it takes the four decisions shoppers actually agonize over and answers each one from the insurers' own policy text, with the exact clause cited: One grounding contract runs through all four: unresolved . The agent never picks a winner. No login needed. The landing page frames the four decisions; each job card and the hero opens the chat modal with its question already prefilled β€” you press Send, it never auto-sends. Try these in order: The landing page frames the four decisions, not a chat box. Each job card opens the modal with its question prefilled. Nothing auto-sends. The answered state: "There's no single winner β€” it depends on when the emergency hits." Verdict first, then the month-by-month math, then the cited clauses. Next.js + TypeScript app. The interesting parts: app/lib/mcp.ts β€” Sanity Context MCP client JSON-RPC over HTTPS . The app never queries Sanity directly; every fact comes through the Context endpoint's tools initial context , groq query , … . app/lib/cost.ts β€” deterministic money math in integer cents: premiums, deductibles, reimbursement rates, annual limits, multi-year scenarios. The LLM explains the numbers; it never computes them. Tested to the cent β€” cost.test.ts . app/lib/insureVsSave.ts , switching.ts , denial.ts , renewals.ts , appeals.ts β€” the four decision jobs as typed modules, each with its own test suite. app/components/ChatModal.tsx + ChatApp.tsx β€” the modal chat: prefilled questions, streaming answers with loading stages, citation cards, contradiction panels, cost cards. The repo is private lymcho/clausehound , now merged to main ; the test suite 1,653 tests and the eval fixtures are part of the submission evidence below. Everything ClauseHound knows lives in Sanity as structured content β€” not blobs of text, but typed documents with relationships, modeled from the insurers' published US sample policies: production dataset, verified live. The schema is the trust strategy. Every clause document carries a required sectionRef e.g. Sec. 3.4 β€” the citation anchor β€” and a required policy reference back to its policyDocument , which references its insurer : insurer β†’ policyDocument β†’ coverageClause / exclusionClause / waitingPeriod / claimStep each with required sectionRef The agent reads through a Sanity Context MCP endpoint backed by an Agent Context over that dataset. A GROQ content filter at the context level scopes every query to the four answerable types and only documents attached to a real policy: type in "coverageClause", "exclusionClause", "waitingPeriod", "claimStep" && defined type == "policyDocument" && id == ^.policy. ref 0 That filter is the guardrail: the agent cannot wander into irrelevant content no matter what the user asks. Coverage, exclusions, and waiting periods are distinct document types and stay distinct in answers β€” which is exactly why the conflict view works: a coverageClause and an exclusionClause matching the same question surface as two typed records with their sectionRef s, not as two paragraphs to blend. The app runs an agentic loop over the MCP tools and narrates the pipeline at real phase boundaries β€” "Pulling the relevant policy text…", "Analyzing your question against the fine print…", "Merging the insurer-by-insurer findings…", "Verifying every claim against the policy text…". The grounding isn't just enforced, it's visible while you wait. This is the bet the challenge asks for: the agent only works because the content was structured. Keyword search over the same policies could find the words "hip dysplasia"; it could not keep five carriers' answers from bleeding into each other, keep a coverage clause distinct from the exclusion that overrides it, or compute a 10-year cost from typed premium/deductible/reimbursement fields. The structure is the product. The release is covered by a 19-fixture held-out eval eval/fixtures.json , fixtures frozen β€” never edited to make the app pass . Each fixture posts a real user question to the deployed /api/chat and checks the response structurally: HTTP 200, citations resolve to real clause IDs, no dangling markers, carrier scope correct, refusals refuse, contradictions surface as unresolved . The two cost fixtures go further: they paste real quotes and re-derive every dollar figure with an independent port of the integer-cent math β€” exact match required, no LLM judging. 16/19 passed on the production build, 2026-10-01 merged to main as 9957cb1 ; the merge tree is byte-identical to the evaluated feature commit 6f4e452b , so the results stand for production . Live-answer latencies ran 18–95s streamed ; the deterministic cost fixtures verify in under a second. Honest failures, since they matter more than the score: Known limitations the eval doesn't hide: yijqzehr , dataset insurer , policyDocument , coverageClause , exclusionClause , waitingPeriod , claimStep SANITY AGENT CONTEXT URL , backed by an Agent Context over the production dataset; the GROQ content filter above scopes all retrieval to answerable, policy-attached documents. No login is needed to try ClauseHound. Live answers are streamed; the deterministic cost comparisons run without the model. Demo only β€” research aid, not professional insurance advice. Verify with your insurer before buying.