{"slug": "auditable-agents-turn-the-answer-into-a-claim-you-can-check", "title": "Auditable agents: turn the answer into a claim you can check", "summary": "A developer released reactifact, an open-source Python framework that makes AI agent answers auditable by storing each output as a typed artifact in a versioned context rather than a message in a log. The approach records typed provenance edges between artifacts and computes a content hash over state, so tracing why an agent produced an answer becomes a graph walk and verifying reproducibility becomes a hash comparison. The package runs offline without an API key, and the author notes the ideas are portable beyond the library.", "body_md": "A model tells you *\"Q2 cloud spend was $45,000 — a 12.5% variance over budget.\"*\n\nNow two uncomfortable questions:\n\nMost agent frameworks can't answer either cleanly. The run is a stream of\n\nmessages; the \"reasoning\" is prose in a log; the numbers came from wherever the\n\nmodel felt like. When something goes wrong in production, your audit trail is a\n\ntranscript you read by eye.\n\nThis is the second post in a short series about building agents whose answers are\n\n**checkable**. The [first one](https://dev.to/bzdvdn/your-ai-agent-re-sends-the-email-on-retry-an-outbox-for-side-effects-13id) was about side effects (an outbox\n\nso a replay doesn't re-send). This one is about the other half: making the *state*\n\nbehind an answer inspectable and reproducible. It's the approach behind\n\n[reactifact](https://github.com/bzdvdn/reactifact), but the ideas — typed\n\nartifacts, typed edges, a content hash over state — are portable.\n\n**TL;DR** — Make the answer a *typed artifact* in a versioned context, not a\n\nmessage in a bag. Link each derivation to its inputs as typed edges. Then \"why\n\ndid it say that?\" is a graph walk and \"is this the same run?\" is a hash — not a\n\ntranscript you read by eye.\n\nRunnable offline in about a minute, no API key (the figures are computed in\n\nPython, nothing calls a model):\n\n```\npip install reactifact\npython -m examples.fintech_audit.main   # prints the audit report, re-hashes to prove it's reproducible\n```\n\nCode: [github.com/bzdvdn/reactifact](https://github.com/bzdvdn/reactifact)\n\n|  | Typical agent framework | reactifact | \n|---|---|---|\n| The answer is | a string in a message log | a typed artifact in the context | \n| State | a message list | versioned commits you can diff | \n| \"Why?\" | read the transcript by eye | walk typed provenance edges | \n| \"Same run?\" | trust the prompt and the model | compare a `context_hash` | \n\nThe core move is boring and it changes everything: **every meaningful thing an agent produces is a typed artifact in one evolving context**, not a message in a\n\n```\nContext v1   Question\nContext v2   + Document, Document\nContext v3   + Evidence, Evidence\nContext v4   + Claim\nContext v5   + VerifiedClaim\nContext v6   + Answer\n```\n\nAn artifact is a pydantic model — `Evidence(text=..., source=...,`, \n\nlocator=\"budget.csv\")`Variance(pct=0.125)`. It has an id, a version, a\n\ncontent hash, and *who produced it*. The context is versioned like a git history:\n\nevery step is a commit you can diff, roll back, check out.\n\nAlready that's more than a transcript. But the interesting part is the edges.\n\nWhen a produce derives something, it doesn't write \"used budget.csv\" into the\n\nprose. It records a **typed relation** — a first-class edge in the artifact\n\ngraph:\n\n```\nsource = self.effects.create(SourceRef(locator=\"transactions.csv\"), id=\"ref:tx\")\ntable = self.effects.create(Table(rows=...), id=\"doc:transactions.csv\")\nspend = self.effects.create(Spend(total=45000.0), id=\"spend:q2\")\nvariance = self.effects.create(Variance(pct=0.125), id=\"variance:q2\")\nanswer = self.effects.create(AuditAnswer(text=\"...$45,000... (+12.5%)\"), id=\"answer:q2\")\n\nanswer.link(\"supported_by\", variance)          # claim ← its calculation\nvariance.link(\"calculated_from\", spend)        # calculation ← its inputs\nspend.link(\"materialized_from\", table)         # figure ← the materialized table\ntable.link(\"materialized_from\", source)        # table ← the source it was read from\n```\n\nThe relations are queryable (`context.related(answer.id, \"supported_by\")`), so\n\n\"why did it say that?\" becomes a graph walk, not a grep. And because the graph is\n\nbuilt, you can render it — Mermaid in the CLI/dashboard, or a structured report:\n\n``` python\nfrom reactifact.audit import build_report, report_to_markdown\n\nreport = build_report(context, answer)   # walks the whole chain, breadth-first\nprint(report_to_markdown(report))\n```\n\n`build_report` returns every artifact that contributed to the answer — each with\n\nits **content hash, version, and producing author** — plus the source locators\n\nthe answer rests on. That's the \"why\": a machine-checkable provenance chain,\n\nnot a paragraph you have to believe.\n\n```\n# Audit report\n- context version: 7\n- context sha256: `f38c6a42…`\n- Answer (sha256 c1d0…, by \"finalize\")\n  - supported_by → Variance (sha256 9a51…, by \"compute_variance\")\n    - calculated_from → Spend (sha256 4f2c…, by \"compute_spend\")\n      - materialized_from → Table (sha256 77b1…, locator \"budget.csv\")\n```\n\nProvenance answers \"why\". **Reproducibility** answers \"is this the same run\".\n\nBecause the context is canonical, you can fingerprint it. `context_hash` is a\n\nsha256 over the run's state — each artifact's id, type, version and content hash,\n\nplus every relation edge. Timestamps are deliberately **excluded**, so two runs\n\nthat reach the same state hash identically:\n\n``` python\nfrom reactifact.audit import context_hash\n\nfirst = context_hash(await run_pipeline())\nsecond = context_hash(await run_pipeline())\nassert first == second         # reproducible — or it fails loudly\n```\n\nThat single string is an audit primitive. Save it next to the answer; later, a\n\nreviewer re-runs the pipeline (or replays a saved session) and compares:\n\n```\nreactifact replay sessions.sqlite3 --session q2 --verify f38c6a42…\n# exits non-zero on any mismatch\n```\n\nNo more \"the numbers look about right.\" The state behind the answer either\n\nhashes to the recorded fingerprint or it doesn't.\n\nA hash only means something if the inputs are controlled. An agent run has three\n\nusual sources of nondeterminism, and each has a handle:\n\n``` python\n  from reactifact.replay import ReplayLLM\n\n  # pass 1 — record a real run\n  resources = RuntimeResources(llm=ReplayLLM(\"calls.jsonl\", mode=\"record\", inner=real_llm))\n  # pass 2 — reproduce it exactly; a divergent call raises ReplayMiss, never guesses\n  resources = RuntimeResources(llm=ReplayLLM(\"calls.jsonl\", mode=\"replay\"))\n```\n\n**Auto-generated ids (`uuid4`) and wall-clock time.** Pass a deterministic id\n\nfactory and stop seeding artifact data from `time.time()` / `uuid4()`. A\n\nrecorded model plus `counter_ids()` is often the whole fix.\n\n**Order and set iteration.** Prefer stable, content-derived ids and explicit\n\nsorting in your produces.\n\nTo catch a leak, run the pipeline a few times under a recorded model and strict\n\nids and compare the fingerprints:\n\n``` python\nfrom reactifact.replay import verify_run\n\nreport = await verify_run(build, recording=\"calls.jsonl\")   # runs it twice\nassert report.ok, report.hashes    # a diff is real nondeterminism in your code\n```\n\nThat's the difference between \"it usually returns the same thing\" and \"a second\n\nrun hashes to the same string.\"\n\nHere's the subtle part, and where this connects to the outbox from the first post.\n\nA naive \"replay\" re-runs the agent — which means it can hit the network again,\n\ncall tools again, and drift. reactifact's replay instead **rebuilds the state from the commit chain without running any agent**:\n\n``` python\nfrom reactifact.replay import replay_context, replay_summary\n\ncontext = await replay_context(store, session_id, version=7)   # state at commit 7\nprint(replay_summary(context))    # counts by artifact type, relations, actions\n```\n\nBecause the commit chain is deterministic, you can reconstruct the exact context\n\nat any point — walk the provenance, render the graph, answer \"why\" — and, since\n\nno agent runs, nothing external fires. Pair it with the outbox and a recorded\n\nside effect is read back as state instead of being re-sent.\n\nThe same versioning makes **alternative states** cheap: `context.branch()` to\n\nexplore two hypotheses, three-way `merge()` with explicit conflicts (no silent\n\nlast-write-wins), `context.diff(v4, v9)` to see exactly what changed between two\n\nturns, `context.checkout(v7)` to move head back and undo a bad step. Time-travel over one artifact\n\ngraph, not a checkpoint of a message list.\n\nOnce the run *is* structured state, evaluation stops being `answer ==`. You can score the layers separately:\n\nexpected_answer\n\n```\nEvidence quality · Claim correctness · Provenance grounding ·\nCalculation correctness · Confidence calibration · Answer quality · Source coverage\n```\n\n`reactifact.eval` runs multi-level metrics over the final `Context` — including\n\nprovenance grounding, i.e. \"is the answer actually linked to evidence that\n\nsupports it?\" — which is a *structural* check, not a model's opinion. And because\n\nthe state is reproducible, a metric that passes today passes on replay.\n\nThe report answers \"why\" from the **final state**. The same design makes the\n\n**runtime trace** worth keeping: a run isn't a wall of log lines you read by\n\neye, it's a directed record of what actually happened. Each agent span carries\n\nthe artifact type that triggered the agent, and which `Produce`(s) ran for that\n\nevent — with how many effect operations each authored and how long it took. So\n\n\"the model said X\" decomposes into \"this event woke this agent, and *this*\n\nproduce did the work\", not a black box.\n\nThat view is built in, not bolted on: the local SQLite dashboard\n\n(`create_trace_router`) shows a `Consume → Produce` flow on each span, and the\n\nsame spans go to Langfuse (one child observation per produce, so the waterfall\n\nshows each step) or any OTLP collector through one `Tracer`. Audit becomes two\n\nviews of one thing — the final provenance graph you can hash, and the causal\n\ntrace that built it.\n\nAuditability here is a property of *your* pipeline, and it only holds as far as\n\nyou make it hold:\n\n`context_hash` excludes timestamps, but if your\nproduce calls `uuid4()` or reads the clock into artifact data, two runs will\ndiffer — `verify_run` tells you, it doesn't fix it.\nThat's the wager: an answer should be a claim you can check, and the machinery to\n\ncheck it — typed artifacts, typed edges, a content hash over state, replay that\n\nreconstructs — is worth building into the framework rather than bolting onto the\n\nlogs afterwards.\n\nThe `fintech_audit` example is exactly the scenario above — a variance over two\n\nCSVs and a policy doc — with no API key (nothing calls a model; the figures are\n\ncomputed in Python):\n\n```\n.venv/bin/python -m examples.fintech_audit.main\n```\n\nIt prints the computed figures, the audit report with a content hash per\n\nartifact, and then runs the pipeline again and asserts the two hashes match — its\n\nexit code is a determinism smoke test.\n\n`docs/en/replay.md`, `docs/en/durability.md`, `docs/en/observability.md`,\n`examples/fintech_audit`\nIf you've shipped agents you had to debug at 2am: **what's your audit trail — logs you read by eye, or state you can query and reproduce?** I'd like to hear", "url": "https://wpnews.pro/news/auditable-agents-turn-the-answer-into-a-claim-you-can-check", "canonical_source": "https://dev.to/bzdvdn/auditable-agents-turn-the-answer-into-a-claim-you-can-check-33i0", "published_at": "2026-10-09 23:32:01+00:00", "updated_at": "2026-10-09 23:58:16.148104+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-tools", "mlops"], "entities": ["reactifact", "pydantic", "GitHub", "Python"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/auditable-agents-turn-the-answer-into-a-claim-you-can-check", "markdown": "https://wpnews.pro/news/auditable-agents-turn-the-answer-into-a-claim-you-can-check.md", "text": "https://wpnews.pro/news/auditable-agents-turn-the-answer-into-a-claim-you-can-check.txt", "jsonld": "https://wpnews.pro/news/auditable-agents-turn-the-answer-into-a-claim-you-can-check.jsonld"}}