Time-Travel Debugging for LLM Agents: Burr's Counterfactual Replay Architecture A developer built Rewind, a counterfactual replay layer on top of Burr, an open-source state machine framework for LLM agent pipelines, that forks a completed run at any step, changes one input, and replays only downstream nodes. The implementation relies on four rules — hashing inputs rather than timestamps, separating tool and LLM caches, replaying only downstream nodes, and diffing by content hash — so unchanged branches return cached outputs at zero token cost. It stores per-node telemetry and state snapshots in SQLite via Burr's SQLitePersister, with a FastAPI backend and React frontend for inspecting runs. Every agent trace tool shows you a waterfall of steps. You spot that step six produced garbage. Now what? You re-run the entire pipeline and hope it lands in the same place. With a non-deterministic model, it doesn't. You can never separate your change from model jitter. The alternative is counterfactual replay: fork a completed run at any step, change exactly one input, replay only the downstream steps, and diff the two trajectories. When you replay a branch you didn't change, every output hash should come back identical and cost zero tokens. That's the difference between a diff you can trust and a diff that's just noise. This is how Rewind works on top of Burr, an open-source state machine framework for agent pipelines. The implementation exposes four design rules that make determinism provable and a handful of traps that break it. Standard logging captures what happened. It doesn't capture why, and it doesn't let you test what would have happened if you changed one decision upstream. What you get from a trace tool: What you can't do: This matters in production when you need to debug a specific run that failed, not reproduce a failure class across ten new runs. The Rewind implementation sits on top of Burr, which already provides a state machine abstraction for agent pipelines. The key additions are persistent state snapshots, per-node telemetry with content hashing, and a replay engine that knows when to use cached outputs. browser Vite + React 18 + TS + Tailwind ├─ Graph · NodeEditor · DiffPanel · CostBar · RawJSON · Settings │ └─ /api Vite proxy ↓ FastAPI spawns daemon thread per run; client polls /graph ↓ Burr application plan → research → analyse → critique → revise → compose → verify → publish llm tool tool llm llm llm tool tool │ ├─ SQLitePersister "burr state" ← state after every node └─ NodeTelemetryHook PostRunStepHook ← inputs/outputs/latency/tokens/hash ↓ rewind.db SQLite, WAL tables: runs · nodes · edges · cache · tool cache · burr state Burr's role: SQLitePersister . Rewind's additions: node name, input hash that stores output hash, output blob, token count . | Rule | Why It Matters | Implementation Detail | |---|---|---| | 1. Hash inputs, not timestamps | Timestamps always change; you need semantic equivalence. | Hash the serialized input dict after stripping metadata keys like timestamp , run id . | | 2. Separate tool cache from LLM cache | Tool calls can be deterministic database query or non-deterministic API with rate limits . | Store tool outputs in tool cache with a TTL or version tag; LLM outputs in cache with no TTL. | | 3. Replay only downstream nodes | Unchanged branches should return cached outputs without re-execution. | Walk the DAG from the fork point; for each node, check if inputs changed. If not, return cached output. | | 4. Diff by content hash, not text | LLM outputs can have whitespace or formatting jitter that doesn't matter. | Store sha256 canonical json output alongside the raw output. Diff hashes first, then show text diff only if hashes differ. | Cost tracking: Simplified replay engine actual implementation in FastAPI route def replay from fork run id: str, fork node: str, new input: dict : original run = load run run id state = load state at node run id, fork node Start from fork point with new input state fork node = new input input hash = hash input new input Check cache cached = query cache fork node, input hash if cached: output = cached "output" tokens = 0 cache hit else: output = execute node fork node, new input tokens = output.get "usage", {} .get "total tokens", 0 store cache fork node, input hash, output, tokens state fork node + " output" = output Walk downstream nodes downstream = get downstream nodes fork node total tokens = tokens for node in downstream: node input = build input from state node, state input hash = hash input node input cached = query cache node, input hash if cached and cached "input hash" == input hash: Unchanged branch: use cache state node + " output" = cached "output" else: Changed branch: execute output = execute node node, node input tokens = output.get "usage", {} .get "total tokens", 0 total tokens += tokens store cache node, input hash, output, tokens state node + " output" = output return { "run id": generate run id , "forked from": run id, "fork node": fork node, "total tokens": total tokens, "state": state } Key points: hash input must be stable: sort dict keys, strip metadata, serialize to canonical JSON. query cache returns None if no match or if input hash differs handles cache invalidation . execute node wraps the Burr node function and extracts token usage from the response. Not all tool calls are deterministic. API rate limits, timestamp drift, and external state changes break replay. Strategies: | Tool Type | Determinism | Cache Strategy | |---|---|---| | Database query read-only | Deterministic if schema stable | Cache indefinitely, keyed by query hash | | External API weather, stock price | Non-deterministic | Cache with TTL e.g., 5 minutes or version tag | | File write | Side effect | Don't cache; log the write and replay with a dry-run flag | | LLM call | Non-deterministic temperature 0 | Cache by input hash; accept that temperature=0 is required for exact replay | Implementation: determinism flag: "deterministic" , "time-bound" , "side-effect" . cached at timestamp and invalidate after TTL. Example: In node definition @action reads= "query" , writes= "results" , determinism="time-bound", ttl=300 def fetch weather state : query = state "query" API call return {"results": call weather api query } During replay, if cached at is older than 300 seconds, re-execute and update cache. The Rewind UI exposes: Metrics tracked per node: input hash : SHA-256 of canonical input JSON. output hash : SHA-256 of canonical output JSON. latency ms : Wall-clock time for node execution. token count : Total tokens prompt + completion for LLM nodes. cache hit : Boolean. Failure modes you can debug: Local development: /api to FastAPI. Production considerations: Scaling: 1. Timestamp leaks into input hash If your node input includes a timestamp or run id , every replay will be a cache miss. Strip metadata before hashing. 2. Non-canonical JSON serialization Python's json.dumps doesn't guarantee key order. Use json.dumps obj, sort keys=True or a library like canonicaljson . 3. LLM temperature 0 Even with identical inputs, the LLM will produce different outputs. Set temperature=0 for deterministic replay, or accept that cache hits only work for exact input matches and you'll need to diff outputs semantically. 4. Tool calls with side effects Writing a file, sending an email, or updating a database breaks replay. Either skip these nodes during replay dry-run mode or log the action without executing. 5. State mutation in node functions If a node mutates shared state e.g., a global cache or config object , replay will see the mutated state from the original run. Ensure node functions are pure or reset shared state before replay. Use Burr + counterfactual replay when: Avoid when: temperature=0 or accept non-deterministic cache misses. Alternatives: The key insight is that deterministic replay requires more than just logging. You need immutable state snapshots, content-addressable caching, and a DAG walker that knows when to skip unchanged branches. Burr provides the state machine primitives; Rewind adds the replay engine and observability layer.