# Time-Travel Debugging for LLM Agents: Burr's Counterfactual Replay Architecture

> Source: <https://dev.to/mech_app_ai/time-travel-debugging-for-llm-agents-burrs-counterfactual-replay-architecture-fno>
> Published: 2026-10-06 00:08:02+00:00

Every agent trace tool shows you a waterfall of steps. You spot that step six produced garbage. Now what? You re-run the entire pipeline and hope it lands in the same place. With a non-deterministic model, it doesn't. You can never separate your change from model jitter.

The alternative is counterfactual replay: fork a completed run at any step, change exactly one input, replay only the downstream steps, and diff the two trajectories. When you replay a branch you didn't change, every output hash should come back identical and cost zero tokens. That's the difference between a diff you can trust and a diff that's just noise.

This is how Rewind works on top of Burr, an open-source state machine framework for agent pipelines. The implementation exposes four design rules that make determinism provable and a handful of traps that break it.

Standard logging captures what happened. It doesn't capture why, and it doesn't let you test what would have happened if you changed one decision upstream.

**What you get from a trace tool:**

**What you can't do:**

This matters in production when you need to debug a specific run that failed, not reproduce a failure class across ten new runs.

The Rewind implementation sits on top of Burr, which already provides a state machine abstraction for agent pipelines. The key additions are persistent state snapshots, per-node telemetry with content hashing, and a replay engine that knows when to use cached outputs.

```
browser (Vite + React 18 + TS + Tailwind)
  ├─ Graph · NodeEditor · DiffPanel · CostBar · RawJSON · Settings
  │
  └─ /api (Vite proxy)
      ↓
    FastAPI (spawns daemon thread per run; client polls /graph)
      ↓
    Burr application
      plan → research → analyse → critique → revise → compose → verify → publish
      (llm)   (tool)     (tool)     (llm)     (llm)     (llm)     (tool)   (tool)
      │
      ├─ SQLitePersister("burr_state") ← state after every node
      └─ NodeTelemetryHook(PostRunStepHook) ← inputs/outputs/latency/tokens/hash
          ↓
        rewind.db (SQLite, WAL)
          tables: runs · nodes · edges · cache · tool_cache · burr_state
```

**Burr's role:**

`SQLitePersister`.
**Rewind's additions:**

`(node_name, input_hash)` that stores `(output_hash, output_blob, token_count)`.
| Rule | Why It Matters | Implementation Detail | 
|---|---|---|
| **1. Hash inputs, not timestamps** | Timestamps always change; you need semantic equivalence. | Hash the serialized input dict after stripping metadata keys like `timestamp` ,`run_id` . | 
| **2. Separate tool cache from LLM cache** | Tool calls can be deterministic (database query) or non-deterministic (API with rate limits). | Store tool outputs in `tool_cache` with a TTL or version tag; LLM outputs in`cache` with no TTL. | 
| **3. Replay only downstream nodes** | Unchanged branches should return cached outputs without re-execution. | Walk the DAG from the fork point; for each node, check if inputs changed. If not, return cached output. | 
| **4. Diff by content hash, not text** | LLM outputs can have whitespace or formatting jitter that doesn't matter. | Store `sha256(canonical_json(output))` alongside the raw output. Diff hashes first, then show text diff only if hashes differ. | 

**Cost tracking:**

```
# Simplified replay engine (actual implementation in FastAPI route)

def replay_from_fork(run_id: str, fork_node: str, new_input: dict):
    original_run = load_run(run_id)
    state = load_state_at_node(run_id, fork_node)

    # Start from fork point with new input
    state[fork_node] = new_input
    input_hash = hash_input(new_input)

    # Check cache
    cached = query_cache(fork_node, input_hash)
    if cached:
        output = cached["output"]
        tokens = 0  # cache hit
    else:
        output = execute_node(fork_node, new_input)
        tokens = output.get("usage", {}).get("total_tokens", 0)
        store_cache(fork_node, input_hash, output, tokens)

    state[fork_node + "_output"] = output

    # Walk downstream nodes
    downstream = get_downstream_nodes(fork_node)
    total_tokens = tokens

    for node in downstream:
        node_input = build_input_from_state(node, state)
        input_hash = hash_input(node_input)

        cached = query_cache(node, input_hash)
        if cached and cached["input_hash"] == input_hash:
            # Unchanged branch: use cache
            state[node + "_output"] = cached["output"]
        else:
            # Changed branch: execute
            output = execute_node(node, node_input)
            tokens = output.get("usage", {}).get("total_tokens", 0)
            total_tokens += tokens
            store_cache(node, input_hash, output, tokens)
            state[node + "_output"] = output

    return {
        "run_id": generate_run_id(),
        "forked_from": run_id,
        "fork_node": fork_node,
        "total_tokens": total_tokens,
        "state": state
    }
```

**Key points:**

`hash_input()` must be stable: sort dict keys, strip metadata, serialize to canonical JSON.`query_cache()` returns `None` if no match or if input hash differs (handles cache invalidation).`execute_node()` wraps the Burr node function and extracts token usage from the response.
Not all tool calls are deterministic. API rate limits, timestamp drift, and external state changes break replay.

**Strategies:**

| Tool Type | Determinism | Cache Strategy | 
|---|---|---|
| Database query (read-only) | Deterministic if schema stable | Cache indefinitely, keyed by query hash | 
| External API (weather, stock price) | Non-deterministic | Cache with TTL (e.g., 5 minutes) or version tag | 
| File write | Side effect | Don't cache; log the write and replay with a dry-run flag | 
| LLM call | Non-deterministic (temperature > 0) | Cache by input hash; accept that temperature=0 is required for exact replay | 

**Implementation:**

`determinism` flag: `"deterministic"`, `"time-bound"`, `"side-effect"`.` cached_at` timestamp and invalidate after TTL.
**Example:**

```
# In node definition
@action(reads=["query"], writes=["results"], determinism="time-bound", ttl=300)
def fetch_weather(state):
    query = state["query"]
    # API call
    return {"results": call_weather_api(query)}
```

During replay, if `cached_at` is older than 300 seconds, re-execute and update cache.

The Rewind UI exposes:

**Metrics tracked per node:**

`input_hash`: SHA-256 of canonical input JSON.` output_hash`: SHA-256 of canonical output JSON.` latency_ms`: Wall-clock time for node execution.` token_count`: Total tokens (prompt + completion) for LLM nodes.` cache_hit`: Boolean.
**Failure modes you can debug:**

**Local development:**

`/api` to FastAPI.
**Production considerations:**

**Scaling:**

**1. Timestamp leaks into input hash**

If your node input includes a `timestamp` or `run_id`, every replay will be a cache miss. Strip metadata before hashing.

**2. Non-canonical JSON serialization**

Python's `json.dumps()` doesn't guarantee key order. Use `json.dumps(obj, sort_keys=True)` or a library like `canonicaljson`.

**3. LLM temperature > 0**

Even with identical inputs, the LLM will produce different outputs. Set `temperature=0` for deterministic replay, or accept that cache hits only work for exact input matches and you'll need to diff outputs semantically.

**4. Tool calls with side effects**

Writing a file, sending an email, or updating a database breaks replay. Either skip these nodes during replay (dry-run mode) or log the action without executing.

**5. State mutation in node functions**

If a node mutates shared state (e.g., a global cache or config object), replay will see the mutated state from the original run. Ensure node functions are pure or reset shared state before replay.

**Use Burr + counterfactual replay when:**

**Avoid when:**

`temperature=0` or accept non-deterministic cache misses.
**Alternatives:**

The key insight is that deterministic replay requires more than just logging. You need immutable state snapshots, content-addressable caching, and a DAG walker that knows when to skip unchanged branches. Burr provides the state machine primitives; Rewind adds the replay engine and observability layer.
