This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange
Fact Ledger β an AI-assisted fact drift detection and remediation engine built on Sanity Content Lake.
Here's the problem it solves, told as a scene:
A product manager at a SaaS company decides to extend refunds from 30 days to 60 days. She opens the CMS, finds the Refund Policy page, updates it, clicks publish. Done. She sends the Slack message: "Refund window is now 60 days, effective immediately."
What she doesn't know: the Help Center article still says 30. The Pricing FAQ still says 30. The Onboarding Guide, the Terms of Service, the Enterprise SLA page, the Checkout confirmation modal copy β all still say 30. Six weeks later, a customer disputes a charge, screenshots the Pricing FAQ, and sends it to their lawyer.
That is Fact Drift β and it happens quietly, constantly, in every company with more than a handful of content pages and more than one person editing them. It's not a CMS problem, it's a structural problem. The number "30" is stored in twenty-three places as dead characters. There's no relationship between them. When one changes, the others don't know.
Fact Ledger fixes this at the data model level.
Business values β refund windows, SLA commitments, file size limits, pricing tiers, data retention periods β become first-class Sanity documents called fact s. Pages don't copy those values; they reference them via a custom factRef inline Portable Text annotation. When the fact document changes, every page that uses a factRef renders the new value instantly, automatically, with zero editor intervention.
For every page that still has the hardcoded plain text β either because it predates the system or because someone didn't know β a deterministic scanner runs on every webhook event, finds the stale copies, and raises structured finding documents. An AI agent drafts the exact Sanity patch mutations to fix them. A human reviews the before/after diffs and clicks one button. Everything updates atomically in a single transaction. The audit ledger captures the whole chain.
The core principle, distilled to four words per step: Rules flag. AI drafts. Human approves. Sanity remembers.
Fact Edit β Webhook β Deterministic Scanner β Finding Graph
β AI Mutation Synthesis β Draft Release β Human Approval β Zero Drift
This is where the design gets interesting. The scanner is pure deterministic TypeScript β no model calls, no embeddings, no fuzzy matching libraries. It runs in milliseconds and produces findings with exact block keys and character offsets.
| Rule | What it catches | Why it matters |
|---|---|---|
| R1 β Unlinked Match | Plain text "30 days" where afactRef should be |
The most common drift; editors hardcode values without realising |
| R2 β Contradiction | Text says "60 days" while the canonical fact says30 |
Catches the pages that were partially updated manually |
| R3 β Deprecated Reference | A factRef pointing to a fact markeddeprecated |
Catches linked pages when a fact is retired, not just changed |
| R4 β Orphan Fact | An active fact exists in Sanity but zero pages reference it | Surfaces forgotten policies that have drifted into irrelevance |
| R5 β Temporal Violation | A factRef pointing to a fact outside itseffectiveFrom/Until range |
Catches seasonal pricing or time-bound policies used out of window |
Run the benchmark against the seeded dataset of 31 planted violations: 100% precision, 100% recall. Every planted issue found, zero false positives. That's not a marketing number β the seed/groundTruth.ts file has the exact expected findings per page, and the benchmark runner verifies against them on every run.
The precision comes from a design decision in R2: the proximity window. Instead of flagging any number that contradicts the fact value anywhere on the page (which would catch phone numbers, copyright years, and table row counts), R2 only flags numbers that appear within N characters of the fact's label in the text. That single detail is the difference between a rule that's useful and one that's noise.
Once findings exist in Sanity, the AI remediation agent kicks in. It fetches the page's Portable Text, locates the exact block and character span identified by the scanner, slices out the stale text, and generates a valid Sanity patch mutation β the kind you'd write by hand if you were doing this manually, except it does it for twenty pages in the time it takes to refresh the dashboard.
All of this goes into a remediation document: a structured release with a fixes array where each item carries beforeText, afterText, the raw patch JSON, and an approved boolean the editor can toggle.
In Sanity Studio, there's a custom "Apply Fixes & Publish" document action on the remediation type. The editor reads through the diffs. They can approve or reject individual fixes. When they're satisfied, one click fires a client.transaction() that atomically:
fixed
changeEvent to the audit ledger
No partial states. No "we fixed 17 of 21 pages and then the tab closed." All twenty-one pages update in the same transaction or none of them do. The Drift Score on the dashboard drops back to zero and stays there β until the next fact changes.
The Next.js dashboard at https://fact-ledger.onrender.com/dashboard is a live, interactive layer on top of Content Lake. It isn't just a read-only frontend β it's a full App SDK Control Center.
1. The Clause Impact Tree
When a policy changes, officials don't just see a list of broken pages. They see a cascading change graph. The Clause Impact Tree visualizes exactly how a single canonical fact ripples outward across the entire corporate document graph, showing nodes turning red in real-time as the webhook fires.
2. The Live Control Center & Charts
Every number and graph on the dashboard comes from a live GROQ query.
3. The Employee Voice Portal
Policy confusion doesn't only flow downward. Employees and customers notice stale policies first. The Employee Voice portal gives that signal a home.
complaint or policyQuestion documents directly in Sanity.
Live Dashboard: https://fact-ledger.onrender.com/dashboard
Sanity Studio: https://pritam.sanity.studio/
Here is how the system actually looks in practice when a fact changes:
1. The Trigger
An editor changes the Cancellation Policy fact value from 60 to 30 in Sanity Studio.
2. The Detection
The webhook fires. The scanner runs R1βR5 across all pages. The dashboard Drift Score jumps and the KPI pulses red.
3. The Remediation Draft
Clicking "Draft AI Remediation" triggers the AI to draft exact before/after patches for all stale spans.
4. The Human Approval
The editor opens the Remediation Release in Studio, reads through the diffs, and clicks "Apply Fixes & Publish".
5. The Resolution
client.transaction() commits. All pages are patched atomically, findings flip to fixed, and the Drift Score returns to 0.
A product manager at a SaaS company decides to extend refunds from 30 days to 60 days. She opens the CMS, updates the Refund Policy page, and clicks publish. What she doesn't know: the Help Center article still says 30. The Pricing FAQ still says 30. The Onboarding Guide, the Terms of Service, the Enterprise SLA page, the Checkout confirmation modal copy β all still say 30.
That is Fact Drift. The number "30" is stored in twenty-three places as dead characters. There's no relationship between them. When one changes, the others don't know.
Fact Ledger fixes this at the data model level. Business values become first-class Sanity documents (facts). Pages don't copy those values; they reference them.
For everyβ¦
fact-ledger/
βββ studio/ # Sanity Studio v3 β 13 schema types, custom structure, Document Action
βββ web/ # Next.js 16 App Router β Drift Dashboard + API routes
βββ scanner/ # Pure TypeScript R1βR5 + Vitest unit tests
βββ seed/ # Idempotent seed script (facts, pages, ground truth)
βββ bench/ # Benchmark runner β writes benchmarkResult docs to Sanity
Sanity Project ID: tmics7hc (dataset: fact-ledger)
I used Antigravity IDE (Google DeepMind) as my primary AI-native IDE, with its browser agent running alongside for live validation and UI feedback during development.
I came into this with a real problem in mind β not a contrived demo. The scenario of "fact changes in one place, stays stale everywhere else" has bitten real teams I've worked with. I wanted to see if a vibe-coded system could actually address it with structural rigor, not just vibes.
My opening prompt was deliberately ambitious:
"Build a closed-loop fact drift detection system on Sanity. Facts are first-class documents. Pages reference them with custom Portable Text inline objects. A deterministic scanner runs on webhook triggers and creates finding documents. An AI agent synthesizes patch mutations. A human approves via a custom Studio document action. Everything is audited. The LLM never publishes anything."
That single prompt produced the full schema architecture, the factRef Portable Text annotation type, and an initial scanner skeleton in one pass. What struck me: the model correctly inferred the key constraint on its own β no unsupervised publishing β and every design decision that followed was shaped by that constraint without me having to repeat it. It understood why the LLM's role should be limited to drafting, not deciding.
Schema design:
"Design a fact schema for Sanity with: key (slug), label, value, unit, aliases (array of strings for surface-form matching), owner (reference to person), dependsOn (self-referential array), effectiveFrom/Until (dates), status (active/deprecated), and a highStakes flag."
The aliases field is what makes the whole system actually useful in practice. Without it, the scanner could only catch exact matches like "30 days". With it, each fact document carries its own list of surface forms: ["30 days", "thirty (30) days", "one month", "30 calendar days"]. R1 matches any of them. That's how you catch the policy writer who wrote "one month" in the terms and the marketing person who wrote "thirty days" in the campaign landing page β same fact, same finding, zero NLP required.
The highStakes flag is a quiet but important detail: facts marked high-stakes require two human approvers in the remediation workflow. One person can't unilaterally ship a fix to the refund policy affecting twenty pages.
Scanner architecture:
"Write the 5 scanner rules as pure TypeScript functions. Zero dependencies. Zero AI. Take facts and Portable Text blocks as input, return findings with exact block key, child key, start offset, and end offset."
The model got genuinely clever here. For R2 (contradiction detection), it implemented a proximity window β a contradictory number only fires as a finding if it appears within N characters of the fact's label in the text body. That design decision alone is what separates a useful rule from a useless one. A naive implementation would flag every number on every page. This one only flags the ones that are contextually near the fact's label, which is where contradictions actually matter.
Running npm run bench against the 31 planted issues: 31 found, 0 false positives. That's a real validation, not a demo.
The custom document action:
"Write a Sanity Studio document action for the remediation type. Label: 'Apply Fixes & Publish'. Filter for approved=true fixes, parse each mutation JSON, build a client.transaction() that patches every page, flips findings to fixed, publishes the remediation doc, writes a changeEvent. Atomic commit."
This came out nearly correct on the first pass. The model chose client.transaction() without prompting β it understood on its own that sequential patches would create a window where some pages are fixed and others aren't. That matters: if the tab closes halfway through twenty patches, you're left with a half-consistent dataset and no way to know which pages were updated. A transaction is all-or-nothing. The model knew that.
Sanity Functions as the webhook handler. The model scaffolded a genuinely clean Sanity Function β right structure, right exports, correct handler signature. Then I went to deploy it and hit the plan wall: Sanity Functions require a paid upgrade. Beautiful code, zero utility at my tier. The fallback was a Next.js API route, which actually turned out cleaner: scanner and webhook in the same process, easier to debug, no cold start latency.
The lesson: beautiful code requiring a plan upgrade isn't code you can ship. We fell back, and the fallback shipped.
Auto-generating seed data. My first two attempts at generating realistic policy pages with planted drift violations produced pages where every violation was placed inside a heading element. Headings in Portable Text don't have child span arrays the way body blocks do β so the scanner couldn't see them. Completely valid Sanity documents, completely invisible to R1βR4.
The fix was to be brutally explicit:
"All planted violations must be inside Portable Text body blocks with at least 2 child spans. Include exact byte offsets in the ground truth file so the benchmark can verify TP/FP/FN."
After that, seed/groundTruth.ts became the source of truth for the benchmark. Every time you run npm run bench, it checks the scanner's output against those exact expected findings. If someone breaks a rule, the benchmark fails β not gracefully, loudly.
The heatmap. Describing the Fact Γ Page heatmap in prose got me a Recharts cell chart with tiles so small you needed a magnifying glass. The breakthrough was switching from describing what it should look like to describing how it should be built:
"Implement this as a CSS grid. Rows = facts, columns = pages. Each cell: red background if open findings exist for that combo, green if clean. No charting library β plain CSS grid. Title attributes for hover tooltips."
One pass, done. The lesson: when the output is wrong, don't describe the outcome differently β describe the implementation.
Workflows as data. The remediation schema is the workflow document. Status transitions: draft β pending_review β approved β published. Each transition is a Sanity patch. Every state is queryable in GROQ. The AI agent advances the document to pending_review. A human editor reviews and advances to published via the custom document action. If they're not ready, the document sits at pending_review indefinitely β visible in the Studio sidebar under "Remediation Releases" with a running count of unapproved fixes. Nothing gets lost. Nothing gets auto-published.
A real-time custom interface. The dashboard is not read-only. Editors can approve or dismiss individual findings directly from the /findings table β the action patches Sanity and the table refreshes with the new state. The Drift Score pulses visually when non-zero. It's a small UX decision, but it matters: you want people to feel the urgency of a non-zero drift score, not just read a number.
The Employee Voice portal. A late addition in Session 3, prompted into existence in a single session. Policy confusion doesn't only flow downward β employees and customers notice it first and report it through informal channels (Slack, email, support tickets) that never get structured or tracked. The portal gives that signal a home in the same Content Lake. A complaint about a confusing refund promise routes directly into a complaint document in Sanity; the editor sees it in Studio next to the open findings for that fact. The problem and its reports live in the same place.
Three schema types, three Studio sidebar sections, five Next.js routes β roughly 45 minutes from prompt to running.
The model never once suggested putting the LLM in the approval loop. Even when I described the remediation flow in ways that left room for autonomous action, it consistently added a canPublish guard, an approved boolean per fix, and explicit before/after text fields for human review. The safety architecture emerged from the model's understanding of the design, not from me specifying every guardrail.
The GROQ queries were also clean on first generation β almost no iteration needed. GROQ's declarative shape maps well onto how the model reasons about data: you describe the shape you want, not the traversal you'd write in SQL. When the output is wrong in SQL, you debug joins. When GROQ output is wrong, you usually just re-describe the shape. The model seems to prefer describing shapes.
The most important moment: the model wrote an elegant Sanity Function. I had to explain that elegance and deployability aren't the same thing. The fallback shipped. The elegant version sits in a comment block as a note for when the plan gets upgraded.
That's the honest build: seven and a half hours, a real problem, a model that got most of it right, a few walls we had to route around, and a system that actually works.
tmics7hc
fact-ledger
fact, factRef (custom Portable Text annotation), page, finding, scanRun, remediation, benchmarkResult, policyUpdate, person
PublishRemediationAction.ts), visionTool for GROQ exploration/api/webhook β scanner β finding creationclient.transaction() for atomic multi-document patchessanity typegen generate in studio/ β all GROQ queries are fully typed
Built with Antigravity IDE (Google DeepMind) over three sessions, roughly 7.5 hours total.
The model wrote the large majority of the TypeScript. My role was architect, reviewer, and the person who hit pricing walls so you don't have to.
Our team came together for the International Charity Day DEV Challenge to build NeedFeed: