{"slug": "i-shipped-an-agent-that-asks-sanity-context-what-it-is-allowed-to-say", "title": "I shipped an agent that asks Sanity Context what it is allowed to say", "summary": "A developer built a travel-assistant agent that is architecturally barred from stating day counts it cannot derive, splitting the work between a deterministic engine that computes every number and a Sanity Context MCP endpoint that supplies the candidate territory set. The classifier returns a choice plus a probability distribution and reports uncertainty below 0.62 confidence, so fragments like \"flew to Paris 20 to 25 February\" are flagged rather than answered fluently. The web app runs on five runtime dependencies with no model-provider SDK.", "body_md": "*This is a submission for the [Sanity Challenge, Path One: Ship an Agent That Queries Real Content](https://dev.to/challenges/sanity-2026-09-16)*\n\n**The agent's MCP endpoint:** `https://api.sanity.io/v1/context/organizations/ovihgdwkx/mcp/ninety`\n\n**Live app:** [https://web-eight-amber-6zft3r0kdf.vercel.app](https://web-eight-amber-6zft3r0kdf.vercel.app)\n\n**Context evidence page:** [https://web-eight-amber-6zft3r0kdf.vercel.app/context](https://web-eight-amber-6zft3r0kdf.vercel.app/context)\n\n**Code:** [https://github.com/PhiBao/beyond-vibe](https://github.com/PhiBao/beyond-vibe)\n\n**Sanity project:** `jvgi63fz` · dataset `production` · Studio: [https://beyond-vibe.sanity.studio](https://beyond-vibe.sanity.studio)\n\nMost travel assistants will confidently tell you how many days you have left. They will also be wrong sometimes, and you cannot tell which times.\n\nSo the agent in this submission is **architecturally forbidden from stating a number.**\n\n```\n// The engine. No model, no network, no env vars.\nevaluate(snapshot, itinerary, {asOf})\n// → {used: 73, attribution: {date: '2026-10-26', blame: [...]}, unresolvedDays: [...]}\n```\n\nThe model may retrieve, classify, explain and cite. Every figure it reports comes from that engine. The split is not stylistic — it is the reason I did not use a language model at all.\n\nType **\"three weeks on Tenerife, then an 8 hour airport layover where I never cleared immigration\"** and it works out what that is, then what it costs.\n\nIt resolves to: Canary Islands, a carve-out, no days counted, 90% confident. Then an airport transit with **no place named**, no days counted, 100% confident. And in between, a leg it is only 50% sure about, so it refuses that one.\n\nThat last behaviour is the whole point. In the middle of that sentence, *\"flew to Paris 20 to 25 February\"* is a fragment, and the agent says so rather than guessing. A chat completion would have answered it fluently and been indistinguishable from a correct answer.\n\n```\ntraveller's words\n     │\n     ├─ date extraction            deterministic, in code, unit-tested\n     │\n     ├─ Sanity Context (MCP)  ──▶  the candidate set. Retrieved, never invented.\n     │                              groq_query over territory + presenceRule\n     │\n     ├─ TypeSafe System One   ──▶  typed classification + calibrated confidence\n     │                              {choice, probabilities, confidence}\n     │\n     └─ the engine            ──▶  every number in the answer\n```\n\nThree properties fall out of that, and each one was bought by a bug I hit while building it:\n\n**1. It cannot invent a place.** The territory list is retrieved from the corpus at request time. If a place is not there, there is no option to select, so the output is `\"No place named\"`. Not a guess at the nearest match — the set does not contain one.\n\n**2. It cannot do arithmetic.** Classification goes through TypeSafe's System One models, which return a value from a set I defined plus the full probability distribution. They are not text generators; there is no channel through which `90` could travel. The whole web app has **five runtime dependencies and no model-provider SDK at all**.\n\n**3. It can say it does not know.** Below 0.62 confidence the read is reported as uncertain rather than answered.\n\nOne consequence worth stating, because it is a bug I shipped and then caught: the classifier also answers *\"would this person be present somewhere that counts against the allowance?\"*, and it does not know that Bulgaria's land crossings were internal in February 2025. So it says no for Sofia-by-train. The engine, which does know, charges ten days. Rather than show two contradicting labels, **the \"N days counted\" figure is read back out of the ledger the engine actually produced** — the classifier never gets to have an opinion about charging.\n\nThe agent's knowledge is a hosted, read-only MCP endpoint. Not a hardcoded prompt, not a vector index built at deploy time — a Sanity-managed endpoint that serves the dataset's schema and answers GROQ on the wire.\n\n| Tool | What the agent uses it for | \n|---|---|\n| `initial_context` | Schema overview: types, fields, relationships, document counts | \n| `groq_query` | Structured retrieval with projections — the candidate set, every request | \n| `schema_explorer` | Field-level detail when the overview is not enough | \n| `array_field_reader` | Reading `accessBands[]` and`competingClaims[]` without pulling whole documents into context | \n\nThe `/context` page calls the endpoint server-side and prints what came back, including its failure modes. Everything on it was fetched when you loaded the page.\n\nReproduce it:\n\n```\ncurl -s https://api.sanity.io/v1/context/organizations/ovihgdwkx/mcp/ninety \\\n  -H \"Authorization: Bearer $SANITY_CONTEXT_TOKEN\" \\\n  -H 'Accept: application/json, text/event-stream' \\\n  -H 'Content-Type: application/json' \\\n  -d '{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/list\"}'\n```\n\nAnd a live `groq_query` through Context, asking about the fact that makes the same fortnight cost fourteen days in Split in 2023 and nothing in 2022:\n\n```\n{\"jsonrpc\":\"2.0\",\"id\":2,\"method\":\"tools/call\",\n \"params\":{\"name\":\"groq_query\",\"arguments\":{\"query\":\n   \"*[_type==\\\"territory\\\"&&code==\\\"HR\\\"][0]{name,\\\"bands\\\":accessBands[]{window{from,to},counted,modes}}\"}}}\n{\"name\":\"Croatia\",\"bands\":[\n  {\"counted\":false,\"window\":{\"from\":\"2000-01-01\",\"to\":\"2023-01-01\"}},\n  {\"counted\":true, \"window\":{\"from\":\"2023-01-01\",\"to\":null}}]}\n```\n\nThat is the product thesis in one payload. A keyword index returns the paragraph about Croatia joining in 2023. This returns the *bands*, which is what you need in order to answer a question about a specific day.\n\nAn agent's answer quality is bounded by the structure it can query. Four decisions do the work.\n\n**Membership is a date-banded array, not a boolean.** `isSchengen: true` is a lie for any country that joined later, and it is the reason most assistants get Croatia and Bulgaria wrong. `accessBands[]` with `{from, to, counted, modes}` makes *\"was this in the area on this day, and by which route\"* answerable at all.\n\n**Carve-outs are territories.** The Canary Islands, Madeira, Åland, Svalbard, the French overseas departments, Ireland, Cyprus and Iceland are each their own `territory` document with their own bands. The agent can therefore say *where* a day was spent, which is also what makes attribution possible.\n\n**Ambiguity is a value, not an exception.** A `presenceRule` can be `counted`, `not_counted`, or `disputed`. When it is `disputed` the engine refuses to classify a day of that kind, and the agent is required to present both readings and name the disagreement. It cannot resolve it, because there is nothing in the corpus to resolve it with.\n\n**Every rule resolves to a source.** Each band, presence rule and regime carries `sourceRef[]` to a real document with a publisher, a URL and a retrieval date. The agent's citations are checkable rather than plausible — and so is the *absence* of one, which is what tells the agent to stop.\n\nThe challenge's own test: *if a keyword search would have gotten you the same answer, aim higher.* So I measured it.\n\nTwo systems, the same corpus, 23 adversarial histories with hand-computed expectations written before the runner existed, so the runner cannot grade itself. **Neither arm uses a language model** — the point is to isolate the contribution of the content structure, and this way anyone can rerun it with `pnpm eval`.\n\n| Measure | Structured (GROQ + engine) | Keyword search (BM25, same rules) | \n|---|---|---|\n| Correct answer | **23/23** | **1/23** | \n| Produced a day count | 21/23 | 0/23 | \n| Named the exact breach date | 1/1 | 0/1 | \n| Refused instead of guessing | 2/2 | n/a | \n\nFull results with the reasoning for each case: `web/eval/RESULTS.md`\n\nThe keyword arm is not handicapped. It is given the same territory records, the same presence rules, the same permit exemptions and the same precedents, flattened into prose, and it scores BM25 over them. It frequently retrieves the right paragraph.\n\nIt cannot produce a verdict, and the reason is worth stating plainly:\n\n**The corpus holds the rules, and never held the traveller.**\n\n\"I have eleven trips logged this year and one booked — am I still legal?\" is not a retrieval problem. It is a join between the rule and eleven date ranges, two of which fall in carve-outs, one of which nobody has ruled on, evaluated day by day across a rolling window. Retrieval finds the paragraphs. It cannot join them to a calendar.\n\nBeing straight about this one: Ninety does **not** use a Sanity Knowledge Base, and the rubric asks for one.\n\nWhat it uses is the dataset served through Context, which covers the structured half of the problem — the rules — but not the prose half. The EU guidance documents the corpus cites are currently referenced as sources, not ingested as searchable text. So an agent question like *\"why is a same-day transit not a stay?\"* can be answered from the structured rule and its citation, but cannot yet quote the underlying paragraph.\n\nIngesting those documents as a Knowledge Base is the next step, and it is the step that would let the agent say *\"the border-crossing page says X, the visa-policy page says Y, and that is why the day is disputed\"* instead of citing two titles and stopping. I would rather name the gap than imply it is closed.\n\n[https://github.com/PhiBao/beyond-vibe](https://github.com/PhiBao/beyond-vibe)\n\n```\nweb/src/lib/agent/describe.ts    free text → dated stays, deterministically\nweb/src/lib/agent/typesafe.ts    typed client for the System One primitives\nweb/src/lib/agent/resolve.ts     retrieve candidates, classify, hand off\nweb/src/lib/sanity/context.ts    JSON-RPC client for the Context MCP endpoint\nweb/src/app/context/page.tsx     the evidence page, rendered server-side\nweb/src/lib/engine/              pure TypeScript. No model, no network.\nweb/scripts/eval/                the 23-case evaluation\n```\n\nA note on why `context.ts` speaks raw JSON-RPC instead of going through an agent SDK: the diagnostics should show what an agent actually receives on the wire, not what a library chooses to show it. Every function degrades to a typed `unavailable` result instead of throwing, so the evidence page reports a missing credential plainly rather than pretending the integration exists.\n\nThere are 56 unit tests. The ones that matter most here are about the free-text layer, because that is where the model is closest to the answer:\n\n`\"1 to 21 February 2026\"` is one stay and not two`\"30 to 31 February 2026\"` is rejected rather than rolled into March\n**Context did not work until the Studio was deployed.** For most of this build the endpoint answered `Only datasets with deployed Studio applications are supported`. It was never an authorisation problem in the end — I had assumed that, and burned time on it. Deploying the Studio fixed it in one command.\n\n**Asking N questions about an array of N stays does not bind question *i* to stay *i*.** I batched all the classifications into one request to save latency. The state was an array of stays and each question said \"this description\", so an 8-hour airport layover in the second stay got classified as the Canary Islands because the first stay mentioned Tenerife. One request per stay now. Cheaper in latency than it sounds, because each request is small.\n\n**`\\s` in a template literal collapses to `s`.** My date-stripping regexes compiled into \"match a literal s\" and silently removed nothing, so the classifier was being handed *\"layover on 3 June 2026\"* with the date still in it. A regex that silently matches nothing is worse than one that throws.\n\n**Two evaluation cases were passing for the wrong reason.** I had written a territory code as `es-canary`, which is a document *key*; the code is `XCI`. The engine's response to an unknown territory is a warning and zero charged days — so the case passed because the Canary days had been dropped rather than classified. The number was right and the reasoning was nonsense. `pnpm check:codes` now fails the build if any case contains a dropped stay.\n\n**I had Iceland in the Schengen Area.** It is EEA, not Schengen, and my corpus said otherwise because I had used a shared helper that stamped the founding date onto every state. Reykjavík days were being charged against 90/180 — a mistake a lot of tools make, made silently, in the one file I was treating as ground truth.\n\nNot legal advice. It covers a handful of passport classes and refuses outside that coverage, which is the only behaviour that makes it safe to trust. 🤖", "url": "https://wpnews.pro/news/i-shipped-an-agent-that-asks-sanity-context-what-it-is-allowed-to-say", "canonical_source": "https://dev.to/kiter/i-shipped-an-agent-that-asks-sanity-context-what-it-is-allowed-to-say-1jol", "published_at": "2026-10-03 18:25:04+00:00", "updated_at": "2026-10-03 18:38:11.533749+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "agent-protocols", "ai-safety"], "entities": ["Sanity", "TypeSafe", "Groq", "Vercel", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-shipped-an-agent-that-asks-sanity-context-what-it-is-allowed-to-say", "markdown": "https://wpnews.pro/news/i-shipped-an-agent-that-asks-sanity-context-what-it-is-allowed-to-say.md", "text": "https://wpnews.pro/news/i-shipped-an-agent-that-asks-sanity-context-what-it-is-allowed-to-say.txt", "jsonld": "https://wpnews.pro/news/i-shipped-an-agent-that-asks-sanity-context-what-it-is-allowed-to-say.jsonld"}}