{"slug": "time-travel-debugging-for-llm-agents-burr-s-counterfactual-replay-architecture", "title": "Time-Travel Debugging for LLM Agents: Burr's Counterfactual Replay Architecture", "summary": "A developer built Rewind, a counterfactual replay layer on top of Burr, an open-source state machine framework for LLM agent pipelines, that forks a completed run at any step, changes one input, and replays only downstream nodes. The implementation relies on four rules — hashing inputs rather than timestamps, separating tool and LLM caches, replaying only downstream nodes, and diffing by content hash — so unchanged branches return cached outputs at zero token cost. It stores per-node telemetry and state snapshots in SQLite via Burr's SQLitePersister, with a FastAPI backend and React frontend for inspecting runs.", "body_md": "Every agent trace tool shows you a waterfall of steps. You spot that step six produced garbage. Now what? You re-run the entire pipeline and hope it lands in the same place. With a non-deterministic model, it doesn't. You can never separate your change from model jitter.\n\nThe alternative is counterfactual replay: fork a completed run at any step, change exactly one input, replay only the downstream steps, and diff the two trajectories. When you replay a branch you didn't change, every output hash should come back identical and cost zero tokens. That's the difference between a diff you can trust and a diff that's just noise.\n\nThis is how Rewind works on top of Burr, an open-source state machine framework for agent pipelines. The implementation exposes four design rules that make determinism provable and a handful of traps that break it.\n\nStandard logging captures what happened. It doesn't capture why, and it doesn't let you test what would have happened if you changed one decision upstream.\n\n**What you get from a trace tool:**\n\n**What you can't do:**\n\nThis matters in production when you need to debug a specific run that failed, not reproduce a failure class across ten new runs.\n\nThe Rewind implementation sits on top of Burr, which already provides a state machine abstraction for agent pipelines. The key additions are persistent state snapshots, per-node telemetry with content hashing, and a replay engine that knows when to use cached outputs.\n\n```\nbrowser (Vite + React 18 + TS + Tailwind)\n  ├─ Graph · NodeEditor · DiffPanel · CostBar · RawJSON · Settings\n  │\n  └─ /api (Vite proxy)\n      ↓\n    FastAPI (spawns daemon thread per run; client polls /graph)\n      ↓\n    Burr application\n      plan → research → analyse → critique → revise → compose → verify → publish\n      (llm)   (tool)     (tool)     (llm)     (llm)     (llm)     (tool)   (tool)\n      │\n      ├─ SQLitePersister(\"burr_state\") ← state after every node\n      └─ NodeTelemetryHook(PostRunStepHook) ← inputs/outputs/latency/tokens/hash\n          ↓\n        rewind.db (SQLite, WAL)\n          tables: runs · nodes · edges · cache · tool_cache · burr_state\n```\n\n**Burr's role:**\n\n`SQLitePersister`.\n**Rewind's additions:**\n\n`(node_name, input_hash)` that stores `(output_hash, output_blob, token_count)`.\n| Rule | Why It Matters | Implementation Detail | \n|---|---|---|\n| **1. Hash inputs, not timestamps** | Timestamps always change; you need semantic equivalence. | Hash the serialized input dict after stripping metadata keys like `timestamp` ,`run_id` . | \n| **2. Separate tool cache from LLM cache** | Tool calls can be deterministic (database query) or non-deterministic (API with rate limits). | Store tool outputs in `tool_cache` with a TTL or version tag; LLM outputs in`cache` with no TTL. | \n| **3. Replay only downstream nodes** | Unchanged branches should return cached outputs without re-execution. | Walk the DAG from the fork point; for each node, check if inputs changed. If not, return cached output. | \n| **4. Diff by content hash, not text** | LLM outputs can have whitespace or formatting jitter that doesn't matter. | Store `sha256(canonical_json(output))` alongside the raw output. Diff hashes first, then show text diff only if hashes differ. | \n\n**Cost tracking:**\n\n```\n# Simplified replay engine (actual implementation in FastAPI route)\n\ndef replay_from_fork(run_id: str, fork_node: str, new_input: dict):\n    original_run = load_run(run_id)\n    state = load_state_at_node(run_id, fork_node)\n\n    # Start from fork point with new input\n    state[fork_node] = new_input\n    input_hash = hash_input(new_input)\n\n    # Check cache\n    cached = query_cache(fork_node, input_hash)\n    if cached:\n        output = cached[\"output\"]\n        tokens = 0  # cache hit\n    else:\n        output = execute_node(fork_node, new_input)\n        tokens = output.get(\"usage\", {}).get(\"total_tokens\", 0)\n        store_cache(fork_node, input_hash, output, tokens)\n\n    state[fork_node + \"_output\"] = output\n\n    # Walk downstream nodes\n    downstream = get_downstream_nodes(fork_node)\n    total_tokens = tokens\n\n    for node in downstream:\n        node_input = build_input_from_state(node, state)\n        input_hash = hash_input(node_input)\n\n        cached = query_cache(node, input_hash)\n        if cached and cached[\"input_hash\"] == input_hash:\n            # Unchanged branch: use cache\n            state[node + \"_output\"] = cached[\"output\"]\n        else:\n            # Changed branch: execute\n            output = execute_node(node, node_input)\n            tokens = output.get(\"usage\", {}).get(\"total_tokens\", 0)\n            total_tokens += tokens\n            store_cache(node, input_hash, output, tokens)\n            state[node + \"_output\"] = output\n\n    return {\n        \"run_id\": generate_run_id(),\n        \"forked_from\": run_id,\n        \"fork_node\": fork_node,\n        \"total_tokens\": total_tokens,\n        \"state\": state\n    }\n```\n\n**Key points:**\n\n`hash_input()` must be stable: sort dict keys, strip metadata, serialize to canonical JSON.`query_cache()` returns `None` if no match or if input hash differs (handles cache invalidation).`execute_node()` wraps the Burr node function and extracts token usage from the response.\nNot all tool calls are deterministic. API rate limits, timestamp drift, and external state changes break replay.\n\n**Strategies:**\n\n| Tool Type | Determinism | Cache Strategy | \n|---|---|---|\n| Database query (read-only) | Deterministic if schema stable | Cache indefinitely, keyed by query hash | \n| External API (weather, stock price) | Non-deterministic | Cache with TTL (e.g., 5 minutes) or version tag | \n| File write | Side effect | Don't cache; log the write and replay with a dry-run flag | \n| LLM call | Non-deterministic (temperature > 0) | Cache by input hash; accept that temperature=0 is required for exact replay | \n\n**Implementation:**\n\n`determinism` flag: `\"deterministic\"`, `\"time-bound\"`, `\"side-effect\"`.` cached_at` timestamp and invalidate after TTL.\n**Example:**\n\n```\n# In node definition\n@action(reads=[\"query\"], writes=[\"results\"], determinism=\"time-bound\", ttl=300)\ndef fetch_weather(state):\n    query = state[\"query\"]\n    # API call\n    return {\"results\": call_weather_api(query)}\n```\n\nDuring replay, if `cached_at` is older than 300 seconds, re-execute and update cache.\n\nThe Rewind UI exposes:\n\n**Metrics tracked per node:**\n\n`input_hash`: SHA-256 of canonical input JSON.` output_hash`: SHA-256 of canonical output JSON.` latency_ms`: Wall-clock time for node execution.` token_count`: Total tokens (prompt + completion) for LLM nodes.` cache_hit`: Boolean.\n**Failure modes you can debug:**\n\n**Local development:**\n\n`/api` to FastAPI.\n**Production considerations:**\n\n**Scaling:**\n\n**1. Timestamp leaks into input hash**\n\nIf your node input includes a `timestamp` or `run_id`, every replay will be a cache miss. Strip metadata before hashing.\n\n**2. Non-canonical JSON serialization**\n\nPython's `json.dumps()` doesn't guarantee key order. Use `json.dumps(obj, sort_keys=True)` or a library like `canonicaljson`.\n\n**3. LLM temperature > 0**\n\nEven with identical inputs, the LLM will produce different outputs. Set `temperature=0` for deterministic replay, or accept that cache hits only work for exact input matches and you'll need to diff outputs semantically.\n\n**4. Tool calls with side effects**\n\nWriting a file, sending an email, or updating a database breaks replay. Either skip these nodes during replay (dry-run mode) or log the action without executing.\n\n**5. State mutation in node functions**\n\nIf a node mutates shared state (e.g., a global cache or config object), replay will see the mutated state from the original run. Ensure node functions are pure or reset shared state before replay.\n\n**Use Burr + counterfactual replay when:**\n\n**Avoid when:**\n\n`temperature=0` or accept non-deterministic cache misses.\n**Alternatives:**\n\nThe key insight is that deterministic replay requires more than just logging. You need immutable state snapshots, content-addressable caching, and a DAG walker that knows when to skip unchanged branches. Burr provides the state machine primitives; Rewind adds the replay engine and observability layer.", "url": "https://wpnews.pro/news/time-travel-debugging-for-llm-agents-burr-s-counterfactual-replay-architecture", "canonical_source": "https://dev.to/mech_app_ai/time-travel-debugging-for-llm-agents-burrs-counterfactual-replay-architecture-fno", "published_at": "2026-10-06 00:08:02+00:00", "updated_at": "2026-10-06 00:17:28.971421+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "large-language-models", "mlops", "ai-tools"], "entities": ["Burr", "Rewind", "SQLitePersister", "FastAPI", "SQLite", "Vite", "React"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/time-travel-debugging-for-llm-agents-burr-s-counterfactual-replay-architecture", "markdown": "https://wpnews.pro/news/time-travel-debugging-for-llm-agents-burr-s-counterfactual-replay-architecture.md", "text": "https://wpnews.pro/news/time-travel-debugging-for-llm-agents-burr-s-counterfactual-replay-architecture.txt", "jsonld": "https://wpnews.pro/news/time-travel-debugging-for-llm-agents-burr-s-counterfactual-replay-architecture.jsonld"}}