cd /news/ai-agents/clausewatch-an-agent-that-refuses-to… · home topics ai-agents article
[ARTICLE · art-134541] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

"ClauseWatch: an agent that refuses to give you one number"

A developer built ClauseWatch, an open-source agent that surfaces conflicts between legal instruments rather than returning a single answer, such as the EU AI Act's six-month log retention floor colliding with the GDPR's unnamed ceiling. The tool pairs a structured Sanity dataset of provisions, requirements, and claims with an LLM agent and a deterministic no-model mode, marking unresolved conflicts as "unmade" with empty decidedBy and decidedAt fields. It runs without credentials against a public query API, with the knowledge base and agent requiring a Context token.

by read11 min views1 publishedSep 19, 2026

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

Ask any model how long you must keep AI system logs under the EU AI Act. It will say six

months. That is true, and it is not the answer.

One instrument sets a floor. Another sets a ceiling. Neither says where the other one sits.

The honest answer is a period you can justify against both — plus a record of who decided it.

ClauseWatch gives that answer. It reads a structured model of obligations and the prose of the

instruments behind them, and when two sources pull in opposite directions it shows both with

their citations and reports that the decision is unmade:

Both clauses stand until someone signs. The dataset keeps a decidedBy and a decidedAt

field for exactly that, and they are empty.

This is the part that needed structured content. A keyword search over the same texts returns

Art. 19 and stops. It cannot know that a clause in a different regulation constrains the

same artifact from the opposite direction, because that fact is not inside either document.

It is in the relationship between them — which is to say, in the schema.

git clone https://github.com/oleg-vdv/clausewatch && cd clausewatch/agent
npm install && npm run web       # http://localhost:4173

No credentials. No login, no token, no API key. The dataset is public, so a fresh clone

reads it over the public query API and renders the full report. The knowledge base and the

LLM agent do need a Context token, which you cannot have — so --check and the page footer

say which of the two sources the answer came from, rather than crediting one they never

touched.

It also runs with no model at all:

npm run ask -- --profile biometric-access --no-llm   # provider, high-risk, KZ → EU
npm run ask -- --profile support-agent --no-llm      # deployer, risk class unset

Two elements of the viewer do the arguing.

The pressure bar — a solid edge on the left where Art. 19 states six months, and a

hatched, edgeless right where the GDPR caps the period without naming one. Drawing a tidy

range there would be a lie about the law.

The signature block — every conflict prints a decided by and a date rule. Resolved,

they carry a name and a date. Open, they are two empty lines in a compliance report. The

thesis as a form field rather than a paragraph.

The Studio is live at https://clausewatch.sanity.studio/, with open conflicts on the first

screen, because that is the one thing an editor of this dataset actually does.

https://github.com/oleg-vdv/clausewatch

studio/   schema, desk structure, seed dataset
agent/    MCP client, domain layer, LLM agent, deterministic report, viewer
demo/     saved runs with their tool calls attached
source        an instrument: binding status, official URL, version label
provision     one citable unit: "Art. 19", verbatim text, effective date, roles bound
requirement   a normalised obligation: "retain automatically generated logs"
  └ claim[]   what each instrument asserts — direction (floor/ceiling/none), period if
              stated, condition, and a reference to its provision
conflict      two clauses that cannot both be satisfied, plus resolution, rationale,
              decidedBy, decidedAt
systemProfile roles, jurisdictions, risk class, agentic or not

Three decisions carry the weight.

Claims sit apart from requirements. One obligation, many instruments, and disagreement

between them is content — not a data-quality problem to be cleaned up.

direction on a claim. An earlier version bucketed claims by does it state a number,

which filed GDPR Art. 5(1)(e) under silence. It is not silent: it caps the period without

naming one. Floor, ceiling and duty-only are three different answers.

A conflict holds a decision, not a resolution rule. An agent that silently picks the

stricter number is guessing on your behalf. One that shows both sides and cites a named,

dated decision produces something an auditor can accept.

Two website sources, both crawled to exact paths rather than wildcards:

Source Paths Documents
artificialintelligenceact.eu Arts. 3, 6, 9, 11–14, 16–21, 26, 27, 72, 73, 99 and Annexes I, III, IV 21
gdpr-info.eu Arts. 5, 17, 25, 30, 32 5

26 documents, indexed into 10 entries. Curated on purpose — and also capped, because the

Context beta allows 150 indexed documents on this plan and I hit that twice before narrowing

properly.

The indexing surprised me. Entries are topic nodes rather than chunks: each carries an article

range, a topic list, cross-references (excludes: … see <other entry>) and a numbered Sources

block with a URL behind every statement.

Two endpoints, because one endpoint serves one mode. Attaching a Knowledge Base to an

endpoint that already served a dataset replaced the GROQ tools entirely.

Endpoint Tools Job
clausewatch-data initial_context ,groq_query ,schema_explorer ,array_field_reader which obligations bind this system
clausewatch-docs initial_context ,knowledge_base_read what the clause says , with its source link

Both carry the same endpoint Instructions — the citation discipline the agent must follow,

delivered before it sees any data. The rule that matters most:

Never fill a gap from your own legal knowledge. If it is not in the dataset or the knowledge

base, say that it is not there.

Beyond the tools I leaned on two Context features: Issues, which surfaced a contradiction

in the corpus, and Instructions, which corrected a fact at the source instead of in my

code. Both stories are below.

A real tool-use loop. Nothing is pre-fetched: the model gets the two endpoints and has to go

and look. A typical run:

1. initial_context         (dataset — schema and the citation rules)
2. kb_initial_context      (knowledge base — the outline of 10 entries)
3. groq_query              (the logging requirement, its claims, its conflicts)
4. kb_knowledge_base_read  (design_documentation, post_market_obligations)
5. groq_query              (jurisdictions, extraterritoriality)
6. kb_knowledge_base_read  (high_risk_categories)
7. groq_query              (system profiles)

Seven calls across both endpoints. The answer states it cannot give one number, shows the

floor and the ceiling with citations and effective dates, reports the conflict as open with

nobody named, and notes the obligation does not bite until 2 December 2027. In a second run

the agent found the matching system profile unprompted, treated riskClass: unknown as

load-bearing, separated provider duties from deployer ones, and closed with a section titled

What I can't answer from these sources.

Saved runs with their tool calls are in

demo/. The answer alone cannot

show that a number came from the endpoints rather than from the model, so the tool calls

travel with it.

The dataset side does the reasoning that a prose search cannot: claims are filtered by the

jurisdictions a system touches and the roles it holds, so a deployer is shown Art. 26(6)

and a provider Art. 19. Telling a deployer that Art. 19 is their duty is not a rounding error,

it is the wrong party.

While indexing, Context flagged that an entry claimed Annex I lists 21 harmonisation

instruments while the source shows 20 — item 1 deleted by amendment, item 21 added. I checked

against the source: true. It also proposed a Section A/B split that I could not verify,

because EUR-Lex blocks automated retrieval.

So the decision recorded in the dataset says the detection was right, states only the verified

part, explicitly excludes the unverified arithmetic, and carries a name and a date.

That is the whole product in one incident. Automated detection is good at finding

disagreement and is not authoritative about resolving it.

In one run the agent cited Art. 26(5) for the deployer's log-retention duty, from the

knowledge base. The source says Art. 26(6); 26(5) is the monitoring duty. A one-digit

error that reads as correct.

Chasing it was the most useful hour of the build. The knowledge-base entry for Article 26

numbers its own sections, and the fifth one is headed ### 5. Log retention. Section five of

an entry about Article 26 — and a model reading it produced "Art. 26(5)". The entry states the

rule itself correctly, and gives no paragraph number at all:

Deployers shall keep automatically generated logs for a period appropriate to the intended

purpose, with a minimum of six months, unless applicable Union or national law … provides

otherwise.

That is the argument for two layers rather than one. Prose is right about the rule and

silent on the address; the dataset carries the address, verified by a human against the

source. Art. 26(6) is now a provision with the checked citation, its own claim on the

retention requirement, and the role it binds — which is why a deployer profile is shown

Art. 26(6) and a provider Art. 19.

I also added a Context Instruction, which is honored on every build, over the raw sources, so a future rebuild cannot reintroduce the wrong number. Being precise about what

The layers check each other in both directions. In another run the agent found a duty present

in the prose that my dataset had not modelled at all, and said so.

Here is the check itself, thirty seconds of it: a script pulls Article 26 from the source and

prints its numbered paragraphs. Paragraph 5 is the duty to monitor. Paragraph 6 is the duty to

keep the logs. That is the whole difference between citing the right clause and the wrong one,

and nothing short of going to the text would have shown it.

Embedded session cards open collapsed — click the ▸ to see the command and everything it printed. Same for the three further down.

count(*) said 31, an anonymous one said 0. Sanity treats _id as a path and public read covers the root path only. An authenticated count is not evidence that your dataset is public./article/* pulled 200 pages including Polish and French translations of the same articles. For a legal agent a translation is a correctness hazard, not noise — and the sitemap is the tell: if a crawl returns more pages than the sitemap lists, it followed in-page language links._source no longer available — and the entry is still listed in initial_context for the agent to read and cite. Dismissing the issue keeps the entry; only a rebuild cleared them.roles[]->name comes back as [{name, _id}]. It also injects _type != "sanity.agentContext" and pins perspective: published. Worth knowing when a query behaves differently in Vision. And one that is nobody's fault but mine: count() of a missing field is null, not 0, so a

role filter written as count(appliesToRoles) == 0 for binds everyone silently dropped every

GDPR claim. It looked fine in the provider's report and only broke for the deployer.

The numbers are the whole story, so here is the terminal output: an import that reports

success, a document fetch that comes back {"documents":[],"omitted":[{"reason":"permission"}]},

a one-document probe with a plain id that proves the ids are the cause, and a final anonymous

count(*) of 18 once everything was renamed to hyphens.

4yzoidsq production (public) Public dataset, no token — every conflict, both sides, with its decision:

https://4yzoidsq.api.sanity.io/v2026-09-19/data/query/production?query=*[_type=="conflict"]{summary,nature,resolution,decidedBy,"sides":sides[]->{"cite":source->shortName+" "+citation}}

The whole build is on record — 485 messages in one Claude Code session. A note before you

open it: I work in Russian, so the conversation is in Russian. Everything that matters here

is not. The tool calls, the GROQ queries, the MCP traces, the terminal output and the agent's

own answers are all English, and that is what each slice below is made of. I have said what

to look for in front of each one — and each card opens collapsed, so click the ▸ to see the

output rather than just the command line.

The question is how long a provider must keep logs for a high-risk biometric system. Watch

the trace: initial_context on the dataset, then on the knowledge base, then a GROQ query it

writes itself, then two knowledge-base entries, then jurisdictions, then profiles. Seven

calls across both Context endpoints before it says anything.

Then read the answer. It gives a floor with a citation, a ceiling with a citation, states

that the conflict between them is unresolved and that nobody is recorded as having decided

it, and points out that the obligation does not apply until December 2027. It never produces

the single number the question was fishing for.

This one starts with a Sanity Context issue: an entry says Annex I lists 21 harmonisation

instruments, the source shows 20. You will see me go to the source rather than take it on

trust — the curl that pulls the annex straight from the source, and the amendment markers

in its output: item 1 struck, item 21 added.

The end of the slice is the part I would point a judge at. Context also proposed a Section

A / Section B split, and I could not verify it, because EUR-Lex blocks automated retrieval.

So the decision recorded in the dataset states the verified part and explicitly excludes the

rest. Finding disagreement and settling it are different jobs.

The knowledge base is built from authoritative reproductions, not the Official Journal —

EUR-Lex blocks automated retrieval. Every source carries an officialUrl and the agent is

instructed to give that one for anything the reader will act on.

The dataset is small and deliberately so: 2 instruments, 9 provisions, 2 requirements,

3 conflicts, 2 system profiles. A working core that demonstrates the model, not a compliance

product. Nothing here is legal advice.

── more in #ai-agents 4 stories · sorted by recency
── more on @clausewatch 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/clausewatch-an-agent…] indexed:0 read:11min 2026-09-19 ·