Auditable agents: turn the answer into a claim you can check A developer released reactifact, an open-source Python framework that makes AI agent answers auditable by storing each output as a typed artifact in a versioned context rather than a message in a log. The approach records typed provenance edges between artifacts and computes a content hash over state, so tracing why an agent produced an answer becomes a graph walk and verifying reproducibility becomes a hash comparison. The package runs offline without an API key, and the author notes the ideas are portable beyond the library. A model tells you "Q2 cloud spend was $45,000 — a 12.5% variance over budget." Now two uncomfortable questions: Most agent frameworks can't answer either cleanly. The run is a stream of messages; the "reasoning" is prose in a log; the numbers came from wherever the model felt like. When something goes wrong in production, your audit trail is a transcript you read by eye. This is the second post in a short series about building agents whose answers are checkable . The first one https://dev.to/bzdvdn/your-ai-agent-re-sends-the-email-on-retry-an-outbox-for-side-effects-13id was about side effects an outbox so a replay doesn't re-send . This one is about the other half: making the state behind an answer inspectable and reproducible. It's the approach behind reactifact https://github.com/bzdvdn/reactifact , but the ideas — typed artifacts, typed edges, a content hash over state — are portable. TL;DR — Make the answer a typed artifact in a versioned context, not a message in a bag. Link each derivation to its inputs as typed edges. Then "why did it say that?" is a graph walk and "is this the same run?" is a hash — not a transcript you read by eye. Runnable offline in about a minute, no API key the figures are computed in Python, nothing calls a model : pip install reactifact python -m examples.fintech audit.main prints the audit report, re-hashes to prove it's reproducible Code: github.com/bzdvdn/reactifact https://github.com/bzdvdn/reactifact | | Typical agent framework | reactifact | |---|---|---| | The answer is | a string in a message log | a typed artifact in the context | | State | a message list | versioned commits you can diff | | "Why?" | read the transcript by eye | walk typed provenance edges | | "Same run?" | trust the prompt and the model | compare a context hash | The core move is boring and it changes everything: every meaningful thing an agent produces is a typed artifact in one evolving context , not a message in a Context v1 Question Context v2 + Document, Document Context v3 + Evidence, Evidence Context v4 + Claim Context v5 + VerifiedClaim Context v6 + Answer An artifact is a pydantic model — Evidence text=..., source=..., , locator="budget.csv" Variance pct=0.125 . It has an id, a version, a content hash, and who produced it . The context is versioned like a git history: every step is a commit you can diff, roll back, check out. Already that's more than a transcript. But the interesting part is the edges. When a produce derives something, it doesn't write "used budget.csv" into the prose. It records a typed relation — a first-class edge in the artifact graph: source = self.effects.create SourceRef locator="transactions.csv" , id="ref:tx" table = self.effects.create Table rows=... , id="doc:transactions.csv" spend = self.effects.create Spend total=45000.0 , id="spend:q2" variance = self.effects.create Variance pct=0.125 , id="variance:q2" answer = self.effects.create AuditAnswer text="...$45,000... +12.5% " , id="answer:q2" answer.link "supported by", variance claim ← its calculation variance.link "calculated from", spend calculation ← its inputs spend.link "materialized from", table figure ← the materialized table table.link "materialized from", source table ← the source it was read from The relations are queryable context.related answer.id, "supported by" , so "why did it say that?" becomes a graph walk, not a grep. And because the graph is built, you can render it — Mermaid in the CLI/dashboard, or a structured report: python from reactifact.audit import build report, report to markdown report = build report context, answer walks the whole chain, breadth-first print report to markdown report build report returns every artifact that contributed to the answer — each with its content hash, version, and producing author — plus the source locators the answer rests on. That's the "why": a machine-checkable provenance chain, not a paragraph you have to believe. Audit report - context version: 7 - context sha256: f38c6a42… - Answer sha256 c1d0…, by "finalize" - supported by → Variance sha256 9a51…, by "compute variance" - calculated from → Spend sha256 4f2c…, by "compute spend" - materialized from → Table sha256 77b1…, locator "budget.csv" Provenance answers "why". Reproducibility answers "is this the same run". Because the context is canonical, you can fingerprint it. context hash is a sha256 over the run's state — each artifact's id, type, version and content hash, plus every relation edge. Timestamps are deliberately excluded , so two runs that reach the same state hash identically: python from reactifact.audit import context hash first = context hash await run pipeline second = context hash await run pipeline assert first == second reproducible — or it fails loudly That single string is an audit primitive. Save it next to the answer; later, a reviewer re-runs the pipeline or replays a saved session and compares: reactifact replay sessions.sqlite3 --session q2 --verify f38c6a42… exits non-zero on any mismatch No more "the numbers look about right." The state behind the answer either hashes to the recorded fingerprint or it doesn't. A hash only means something if the inputs are controlled. An agent run has three usual sources of nondeterminism, and each has a handle: python from reactifact.replay import ReplayLLM pass 1 — record a real run resources = RuntimeResources llm=ReplayLLM "calls.jsonl", mode="record", inner=real llm pass 2 — reproduce it exactly; a divergent call raises ReplayMiss, never guesses resources = RuntimeResources llm=ReplayLLM "calls.jsonl", mode="replay" Auto-generated ids uuid4 and wall-clock time. Pass a deterministic id factory and stop seeding artifact data from time.time / uuid4 . A recorded model plus counter ids is often the whole fix. Order and set iteration. Prefer stable, content-derived ids and explicit sorting in your produces. To catch a leak, run the pipeline a few times under a recorded model and strict ids and compare the fingerprints: python from reactifact.replay import verify run report = await verify run build, recording="calls.jsonl" runs it twice assert report.ok, report.hashes a diff is real nondeterminism in your code That's the difference between "it usually returns the same thing" and "a second run hashes to the same string." Here's the subtle part, and where this connects to the outbox from the first post. A naive "replay" re-runs the agent — which means it can hit the network again, call tools again, and drift. reactifact's replay instead rebuilds the state from the commit chain without running any agent : python from reactifact.replay import replay context, replay summary context = await replay context store, session id, version=7 state at commit 7 print replay summary context counts by artifact type, relations, actions Because the commit chain is deterministic, you can reconstruct the exact context at any point — walk the provenance, render the graph, answer "why" — and, since no agent runs, nothing external fires. Pair it with the outbox and a recorded side effect is read back as state instead of being re-sent. The same versioning makes alternative states cheap: context.branch to explore two hypotheses, three-way merge with explicit conflicts no silent last-write-wins , context.diff v4, v9 to see exactly what changed between two turns, context.checkout v7 to move head back and undo a bad step. Time-travel over one artifact graph, not a checkpoint of a message list. Once the run is structured state, evaluation stops being answer == . You can score the layers separately: expected answer Evidence quality · Claim correctness · Provenance grounding · Calculation correctness · Confidence calibration · Answer quality · Source coverage reactifact.eval runs multi-level metrics over the final Context — including provenance grounding, i.e. "is the answer actually linked to evidence that supports it?" — which is a structural check, not a model's opinion. And because the state is reproducible, a metric that passes today passes on replay. The report answers "why" from the final state . The same design makes the runtime trace worth keeping: a run isn't a wall of log lines you read by eye, it's a directed record of what actually happened. Each agent span carries the artifact type that triggered the agent, and which Produce s ran for that event — with how many effect operations each authored and how long it took. So "the model said X" decomposes into "this event woke this agent, and this produce did the work", not a black box. That view is built in, not bolted on: the local SQLite dashboard create trace router shows a Consume → Produce flow on each span, and the same spans go to Langfuse one child observation per produce, so the waterfall shows each step or any OTLP collector through one Tracer . Audit becomes two views of one thing — the final provenance graph you can hash, and the causal trace that built it. Auditability here is a property of your pipeline, and it only holds as far as you make it hold: context hash excludes timestamps, but if your produce calls uuid4 or reads the clock into artifact data, two runs will differ — verify run tells you, it doesn't fix it. That's the wager: an answer should be a claim you can check, and the machinery to check it — typed artifacts, typed edges, a content hash over state, replay that reconstructs — is worth building into the framework rather than bolting onto the logs afterwards. The fintech audit example is exactly the scenario above — a variance over two CSVs and a policy doc — with no API key nothing calls a model; the figures are computed in Python : .venv/bin/python -m examples.fintech audit.main It prints the computed figures, the audit report with a content hash per artifact, and then runs the pipeline again and asserts the two hashes match — its exit code is a determinism smoke test. docs/en/replay.md , docs/en/durability.md , docs/en/observability.md , examples/fintech audit If you've shipped agents you had to debug at 2am: what's your audit trail — logs you read by eye, or state you can query and reproduce? I'd like to hear