Your Policies Are Out of Date: How I Built a Sanity AI Agent to Catch Fact Drift A developer built Fact Ledger, an AI-assisted fact drift detection and remediation engine on Sanity Content Lake that stores business values such as refund windows and SLA commitments as first-class "fact" documents referenced by pages via a custom factRef annotation. A deterministic TypeScript scanner flags stale hardcoded copies across five rule types, an AI agent drafts the Sanity patch mutations, and a human approves the diffs before they are applied atomically. The developer reports the scanner achieved 100% precision and 100% recall against a seeded dataset of 31 planted violations. This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange https://dev.to/challenges/sanity-2026-09-16 Fact Ledger — an AI-assisted fact drift detection and remediation engine built on Sanity Content Lake. Here's the problem it solves, told as a scene: A product manager at a SaaS company decides to extend refunds from 30 days to 60 days. She opens the CMS, finds the Refund Policy page, updates it, clicks publish. Done. She sends the Slack message: "Refund window is now 60 days, effective immediately." What she doesn't know: the Help Center article still says 30. The Pricing FAQ still says 30. The Onboarding Guide, the Terms of Service, the Enterprise SLA page, the Checkout confirmation modal copy — all still say 30. Six weeks later, a customer disputes a charge, screenshots the Pricing FAQ, and sends it to their lawyer. That is Fact Drift — and it happens quietly, constantly, in every company with more than a handful of content pages and more than one person editing them. It's not a CMS problem, it's a structural problem. The number "30" is stored in twenty-three places as dead characters. There's no relationship between them. When one changes, the others don't know. Fact Ledger fixes this at the data model level. Business values — refund windows, SLA commitments, file size limits, pricing tiers, data retention periods — become first-class Sanity documents called fact s. Pages don't copy those values; they reference them via a custom factRef inline Portable Text annotation. When the fact document changes, every page that uses a factRef renders the new value instantly, automatically, with zero editor intervention. For every page that still has the hardcoded plain text — either because it predates the system or because someone didn't know — a deterministic scanner runs on every webhook event, finds the stale copies, and raises structured finding documents. An AI agent drafts the exact Sanity patch mutations to fix them. A human reviews the before/after diffs and clicks one button. Everything updates atomically in a single transaction. The audit ledger captures the whole chain. The core principle, distilled to four words per step: Rules flag. AI drafts. Human approves. Sanity remembers. Fact Edit → Webhook → Deterministic Scanner → Finding Graph → AI Mutation Synthesis → Draft Release → Human Approval → Zero Drift This is where the design gets interesting. The scanner is pure deterministic TypeScript — no model calls, no embeddings, no fuzzy matching libraries. It runs in milliseconds and produces findings with exact block keys and character offsets. | Rule | What it catches | Why it matters | |---|---|---| | R1 — Unlinked Match | Plain text "30 days" where a factRef should be | The most common drift; editors hardcode values without realising | | R2 — Contradiction | Text says "60 days" while the canonical fact says 30 | Catches the pages that were partially updated manually | | R3 — Deprecated Reference | A factRef pointing to a fact marked deprecated | Catches linked pages when a fact is retired, not just changed | | R4 — Orphan Fact | An active fact exists in Sanity but zero pages reference it | Surfaces forgotten policies that have drifted into irrelevance | | R5 — Temporal Violation | A factRef pointing to a fact outside its effectiveFrom/Until range | Catches seasonal pricing or time-bound policies used out of window | Run the benchmark against the seeded dataset of 31 planted violations: 100% precision, 100% recall . Every planted issue found, zero false positives. That's not a marketing number — the seed/groundTruth.ts file has the exact expected findings per page, and the benchmark runner verifies against them on every run. The precision comes from a design decision in R2: the proximity window . Instead of flagging any number that contradicts the fact value anywhere on the page which would catch phone numbers, copyright years, and table row counts , R2 only flags numbers that appear within N characters of the fact's label in the text. That single detail is the difference between a rule that's useful and one that's noise. Once findings exist in Sanity, the AI remediation agent kicks in. It fetches the page's Portable Text, locates the exact block and character span identified by the scanner, slices out the stale text, and generates a valid Sanity patch mutation — the kind you'd write by hand if you were doing this manually, except it does it for twenty pages in the time it takes to refresh the dashboard. All of this goes into a remediation document: a structured release with a fixes array where each item carries beforeText , afterText , the raw patch JSON, and an approved boolean the editor can toggle. In Sanity Studio, there's a custom "Apply Fixes & Publish" document action on the remediation type. The editor reads through the diffs. They can approve or reject individual fixes. When they're satisfied, one click fires a client.transaction that atomically: fixed changeEvent to the audit ledger No partial states. No "we fixed 17 of 21 pages and then the tab closed." All twenty-one pages update in the same transaction or none of them do. The Drift Score on the dashboard drops back to zero and stays there — until the next fact changes. The Next.js dashboard at https://fact-ledger.onrender.com/dashboard is a live, interactive layer on top of Content Lake. It isn't just a read-only frontend — it's a full App SDK Control Center . 1. The Clause Impact Tree When a policy changes, officials don't just see a list of broken pages. They see a cascading change graph . The Clause Impact Tree visualizes exactly how a single canonical fact ripples outward across the entire corporate document graph, showing nodes turning red in real-time as the webhook fires. 2. The Live Control Center & Charts Every number and graph on the dashboard comes from a live GROQ query. 3. The Employee Voice Portal Policy confusion doesn't only flow downward. Employees and customers notice stale policies first. The Employee Voice portal gives that signal a home. complaint or policyQuestion documents directly in Sanity. Live Dashboard: https://fact-ledger.onrender.com/dashboard Sanity Studio: https://pritam.sanity.studio/ Here is how the system actually looks in practice when a fact changes: 1. The Trigger An editor changes the Cancellation Policy fact value from 60 to 30 in Sanity Studio. 2. The Detection The webhook fires. The scanner runs R1–R5 across all pages. The dashboard Drift Score jumps and the KPI pulses red. 3. The Remediation Draft Clicking "Draft AI Remediation" triggers the AI to draft exact before/after patches for all stale spans. 4. The Human Approval The editor opens the Remediation Release in Studio, reads through the diffs, and clicks "Apply Fixes & Publish". 5. The Resolution client.transaction commits. All pages are patched atomically, findings flip to fixed, and the Drift Score returns to 0. A product manager at a SaaS company decides to extend refunds from 30 days to 60 days. She opens the CMS, updates the Refund Policy page, and clicks publish. What she doesn't know: the Help Center article still says 30. The Pricing FAQ still says 30. The Onboarding Guide, the Terms of Service, the Enterprise SLA page, the Checkout confirmation modal copy — all still say 30. That is Fact Drift. The number "30" is stored in twenty-three places as dead characters. There's no relationship between them. When one changes, the others don't know. Fact Ledger fixes this at the data model level. Business values become first-class Sanity documents facts . Pages don't copy those values; they reference them. For every… fact-ledger/ ├── studio/ Sanity Studio v3 — 13 schema types, custom structure, Document Action ├── web/ Next.js 16 App Router — Drift Dashboard + API routes ├── scanner/ Pure TypeScript R1–R5 + Vitest unit tests ├── seed/ Idempotent seed script facts, pages, ground truth └── bench/ Benchmark runner — writes benchmarkResult docs to Sanity Sanity Project ID: tmics7hc dataset: fact-ledger I used Antigravity IDE Google DeepMind as my primary AI-native IDE, with its browser agent running alongside for live validation and UI feedback during development. I came into this with a real problem in mind — not a contrived demo. The scenario of "fact changes in one place, stays stale everywhere else" has bitten real teams I've worked with. I wanted to see if a vibe-coded system could actually address it with structural rigor, not just vibes. My opening prompt was deliberately ambitious: "Build a closed-loop fact drift detection system on Sanity. Facts are first-class documents. Pages reference them with custom Portable Text inline objects. A deterministic scanner runs on webhook triggers and creates finding documents. An AI agent synthesizes patch mutations. A human approves via a custom Studio document action. Everything is audited. The LLM never publishes anything." That single prompt produced the full schema architecture, the factRef Portable Text annotation type, and an initial scanner skeleton in one pass. What struck me: the model correctly inferred the key constraint on its own — no unsupervised publishing — and every design decision that followed was shaped by that constraint without me having to repeat it. It understood why the LLM's role should be limited to drafting, not deciding. Schema design: "Design a fact schema for Sanity with: key slug , label, value, unit, aliases array of strings for surface-form matching , owner reference to person , dependsOn self-referential array , effectiveFrom/Until dates , status active/deprecated , and a highStakes flag." The aliases field is what makes the whole system actually useful in practice. Without it, the scanner could only catch exact matches like "30 days" . With it, each fact document carries its own list of surface forms: "30 days", "thirty 30 days", "one month", "30 calendar days" . R1 matches any of them. That's how you catch the policy writer who wrote "one month" in the terms and the marketing person who wrote "thirty days" in the campaign landing page — same fact, same finding, zero NLP required. The highStakes flag is a quiet but important detail: facts marked high-stakes require two human approvers in the remediation workflow. One person can't unilaterally ship a fix to the refund policy affecting twenty pages. Scanner architecture: "Write the 5 scanner rules as pure TypeScript functions. Zero dependencies. Zero AI. Take facts and Portable Text blocks as input, return findings with exact block key, child key, start offset, and end offset." The model got genuinely clever here. For R2 contradiction detection , it implemented a proximity window — a contradictory number only fires as a finding if it appears within N characters of the fact's label in the text body. That design decision alone is what separates a useful rule from a useless one. A naive implementation would flag every number on every page. This one only flags the ones that are contextually near the fact's label, which is where contradictions actually matter. Running npm run bench against the 31 planted issues: 31 found, 0 false positives . That's a real validation, not a demo. The custom document action: "Write a Sanity Studio document action for the remediation type. Label: 'Apply Fixes & Publish'. Filter for approved=true fixes, parse each mutation JSON, build a client.transaction that patches every page, flips findings to fixed, publishes the remediation doc, writes a changeEvent. Atomic commit." This came out nearly correct on the first pass. The model chose client.transaction without prompting — it understood on its own that sequential patches would create a window where some pages are fixed and others aren't. That matters: if the tab closes halfway through twenty patches, you're left with a half-consistent dataset and no way to know which pages were updated. A transaction is all-or-nothing. The model knew that. Sanity Functions as the webhook handler. The model scaffolded a genuinely clean Sanity Function — right structure, right exports, correct handler signature. Then I went to deploy it and hit the plan wall: Sanity Functions require a paid upgrade. Beautiful code, zero utility at my tier. The fallback was a Next.js API route, which actually turned out cleaner: scanner and webhook in the same process, easier to debug, no cold start latency. The lesson: beautiful code requiring a plan upgrade isn't code you can ship. We fell back, and the fallback shipped. Auto-generating seed data. My first two attempts at generating realistic policy pages with planted drift violations produced pages where every violation was placed inside a heading element. Headings in Portable Text don't have child span arrays the way body blocks do — so the scanner couldn't see them. Completely valid Sanity documents, completely invisible to R1–R4. The fix was to be brutally explicit: "All planted violations must be inside Portable Text body blocks with at least 2 child spans. Include exact byte offsets in the ground truth file so the benchmark can verify TP/FP/FN." After that, seed/groundTruth.ts became the source of truth for the benchmark. Every time you run npm run bench , it checks the scanner's output against those exact expected findings. If someone breaks a rule, the benchmark fails — not gracefully, loudly. The heatmap. Describing the Fact × Page heatmap in prose got me a Recharts cell chart with tiles so small you needed a magnifying glass. The breakthrough was switching from describing what it should look like to describing how it should be built : "Implement this as a CSS grid. Rows = facts, columns = pages. Each cell: red background if open findings exist for that combo, green if clean. No charting library — plain CSS grid. Title attributes for hover tooltips." One pass, done. The lesson: when the output is wrong, don't describe the outcome differently — describe the implementation. Workflows as data. The remediation schema is the workflow document. Status transitions: draft → pending review → approved → published . Each transition is a Sanity patch. Every state is queryable in GROQ. The AI agent advances the document to pending review . A human editor reviews and advances to published via the custom document action. If they're not ready, the document sits at pending review indefinitely — visible in the Studio sidebar under "Remediation Releases" with a running count of unapproved fixes. Nothing gets lost. Nothing gets auto-published. A real-time custom interface. The dashboard is not read-only. Editors can approve or dismiss individual findings directly from the /findings table — the action patches Sanity and the table refreshes with the new state. The Drift Score pulses visually when non-zero. It's a small UX decision, but it matters: you want people to feel the urgency of a non-zero drift score, not just read a number. The Employee Voice portal. A late addition in Session 3, prompted into existence in a single session. Policy confusion doesn't only flow downward — employees and customers notice it first and report it through informal channels Slack, email, support tickets that never get structured or tracked. The portal gives that signal a home in the same Content Lake. A complaint about a confusing refund promise routes directly into a complaint document in Sanity; the editor sees it in Studio next to the open findings for that fact. The problem and its reports live in the same place. Three schema types, three Studio sidebar sections, five Next.js routes — roughly 45 minutes from prompt to running. The model never once suggested putting the LLM in the approval loop. Even when I described the remediation flow in ways that left room for autonomous action, it consistently added a canPublish guard, an approved boolean per fix, and explicit before/after text fields for human review. The safety architecture emerged from the model's understanding of the design, not from me specifying every guardrail. The GROQ queries were also clean on first generation — almost no iteration needed. GROQ's declarative shape maps well onto how the model reasons about data: you describe the shape you want, not the traversal you'd write in SQL. When the output is wrong in SQL, you debug joins. When GROQ output is wrong, you usually just re-describe the shape. The model seems to prefer describing shapes. The most important moment: the model wrote an elegant Sanity Function. I had to explain that elegance and deployability aren't the same thing. The fallback shipped. The elegant version sits in a comment block as a note for when the plan gets upgraded. That's the honest build: seven and a half hours, a real problem, a model that got most of it right, a few walls we had to route around, and a system that actually works. tmics7hc fact-ledger fact , factRef custom Portable Text annotation , page , finding , scanRun , remediation , benchmarkResult , policyUpdate , person PublishRemediationAction.ts , visionTool for GROQ exploration /api/webhook → scanner → finding creation client.transaction for atomic multi-document patches sanity typegen generate in studio/ — all GROQ queries are fully typed Built with Antigravity IDE Google DeepMind over three sessions, roughly 7.5 hours total. The model wrote the large majority of the TypeScript. My role was architect, reviewer, and the person who hit pricing walls so you don't have to. Our team came together for the International Charity Day DEV Challenge to build NeedFeed: