cd /news/ai-agents/time-travel-debugging-for-llm-agents… · home › topics › ai-agents › article
[ARTICLE · art-145748] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Time-Travel Debugging for LLM Agents: Burr's Counterfactual Replay Architecture

A developer built Rewind, a counterfactual replay layer on top of Burr, an open-source state machine framework for LLM agent pipelines, that forks a completed run at any step, changes one input, and replays only downstream nodes. The implementation relies on four rules — hashing inputs rather than timestamps, separating tool and LLM caches, replaying only downstream nodes, and diffing by content hash — so unchanged branches return cached outputs at zero token cost. It stores per-node telemetry and state snapshots in SQLite via Burr's SQLitePersister, with a FastAPI backend and React frontend for inspecting runs.

by read6 min views2 publishedOct 6, 2026

Every agent trace tool shows you a waterfall of steps. You spot that step six produced garbage. Now what? You re-run the entire pipeline and hope it lands in the same place. With a non-deterministic model, it doesn't. You can never separate your change from model jitter.

The alternative is counterfactual replay: fork a completed run at any step, change exactly one input, replay only the downstream steps, and diff the two trajectories. When you replay a branch you didn't change, every output hash should come back identical and cost zero tokens. That's the difference between a diff you can trust and a diff that's just noise.

This is how Rewind works on top of Burr, an open-source state machine framework for agent pipelines. The implementation exposes four design rules that make determinism provable and a handful of traps that break it.

Standard logging captures what happened. It doesn't capture why, and it doesn't let you test what would have happened if you changed one decision upstream.

What you get from a trace tool:

What you can't do:

This matters in production when you need to debug a specific run that failed, not reproduce a failure class across ten new runs.

The Rewind implementation sits on top of Burr, which already provides a state machine abstraction for agent pipelines. The key additions are persistent state snapshots, per-node telemetry with content hashing, and a replay engine that knows when to use cached outputs.

browser (Vite + React 18 + TS + Tailwind)
  ├─ Graph · NodeEditor · DiffPanel · CostBar · RawJSON · Settings
  │
  └─ /api (Vite proxy)
      ↓
    FastAPI (spawns daemon thread per run; client polls /graph)
      ↓
    Burr application
      plan → research → analyse → critique → revise → compose → verify → publish
      (llm)   (tool)     (tool)     (llm)     (llm)     (llm)     (tool)   (tool)
      │
      ├─ SQLitePersister("burr_state") ← state after every node
      └─ NodeTelemetryHook(PostRunStepHook) ← inputs/outputs/latency/tokens/hash
          ↓
        rewind.db (SQLite, WAL)
          tables: runs · nodes · edges · cache · tool_cache · burr_state

Burr's role:

SQLitePersister. Rewind's additions:

(node_name, input_hash) that stores (output_hash, output_blob, token_count).

Rule Why It Matters Implementation Detail
1. Hash inputs, not timestamps Timestamps always change; you need semantic equivalence. Hash the serialized input dict after stripping metadata keys like timestamp ,run_id .
2. Separate tool cache from LLM cache Tool calls can be deterministic (database query) or non-deterministic (API with rate limits). Store tool outputs in tool_cache with a TTL or version tag; LLM outputs incache with no TTL.
3. Replay only downstream nodes Unchanged branches should return cached outputs without re-execution. Walk the DAG from the fork point; for each node, check if inputs changed. If not, return cached output.
4. Diff by content hash, not text LLM outputs can have whitespace or formatting jitter that doesn't matter. Store sha256(canonical_json(output)) alongside the raw output. Diff hashes first, then show text diff only if hashes differ.

Cost tracking:


def replay_from_fork(run_id: str, fork_node: str, new_input: dict):
    original_run = load_run(run_id)
    state = load_state_at_node(run_id, fork_node)

    state[fork_node] = new_input
    input_hash = hash_input(new_input)

    cached = query_cache(fork_node, input_hash)
    if cached:
        output = cached["output"]
        tokens = 0  # cache hit
    else:
        output = execute_node(fork_node, new_input)
        tokens = output.get("usage", {}).get("total_tokens", 0)
        store_cache(fork_node, input_hash, output, tokens)

    state[fork_node + "_output"] = output

    downstream = get_downstream_nodes(fork_node)
    total_tokens = tokens

    for node in downstream:
        node_input = build_input_from_state(node, state)
        input_hash = hash_input(node_input)

        cached = query_cache(node, input_hash)
        if cached and cached["input_hash"] == input_hash:
            state[node + "_output"] = cached["output"]
        else:
            output = execute_node(node, node_input)
            tokens = output.get("usage", {}).get("total_tokens", 0)
            total_tokens += tokens
            store_cache(node, input_hash, output, tokens)
            state[node + "_output"] = output

    return {
        "run_id": generate_run_id(),
        "forked_from": run_id,
        "fork_node": fork_node,
        "total_tokens": total_tokens,
        "state": state
    }

Key points:

hash_input() must be stable: sort dict keys, strip metadata, serialize to canonical JSON.query_cache() returns None if no match or if input hash differs (handles cache invalidation).execute_node() wraps the Burr node function and extracts token usage from the response. Not all tool calls are deterministic. API rate limits, timestamp drift, and external state changes break replay.

Strategies:

Tool Type Determinism Cache Strategy
Database query (read-only) Deterministic if schema stable Cache indefinitely, keyed by query hash
External API (weather, stock price) Non-deterministic Cache with TTL (e.g., 5 minutes) or version tag
File write Side effect Don't cache; log the write and replay with a dry-run flag
LLM call Non-deterministic (temperature > 0) Cache by input hash; accept that temperature=0 is required for exact replay

Implementation:

determinism flag: "deterministic", "time-bound", "side-effect". cached_at timestamp and invalidate after TTL. Example:

@action(reads=["query"], writes=["results"], determinism="time-bound", ttl=300)
def fetch_weather(state):
    query = state["query"]
    return {"results": call_weather_api(query)}

During replay, if cached_at is older than 300 seconds, re-execute and update cache.

The Rewind UI exposes:

Metrics tracked per node:

input_hash: SHA-256 of canonical input JSON. output_hash: SHA-256 of canonical output JSON. latency_ms: Wall-clock time for node execution. token_count: Total tokens (prompt + completion) for LLM nodes. cache_hit: Boolean. Failure modes you can debug:

Local development:

/api to FastAPI. Production considerations:

Scaling:

1. Timestamp leaks into input hash

If your node input includes a timestamp or run_id, every replay will be a cache miss. Strip metadata before hashing.

2. Non-canonical JSON serialization

Python's json.dumps() doesn't guarantee key order. Use json.dumps(obj, sort_keys=True) or a library like canonicaljson.

3. LLM temperature > 0

Even with identical inputs, the LLM will produce different outputs. Set temperature=0 for deterministic replay, or accept that cache hits only work for exact input matches and you'll need to diff outputs semantically.

4. Tool calls with side effects

Writing a file, sending an email, or updating a database breaks replay. Either skip these nodes during replay (dry-run mode) or log the action without executing.

5. State mutation in node functions

If a node mutates shared state (e.g., a global cache or config object), replay will see the mutated state from the original run. Ensure node functions are pure or reset shared state before replay.

Use Burr + counterfactual replay when:

Avoid when:

temperature=0 or accept non-deterministic cache misses. Alternatives:

The key insight is that deterministic replay requires more than just logging. You need immutable state snapshots, content-addressable caching, and a DAG walker that knows when to skip unchanged branches. Burr provides the state machine primitives; Rewind adds the replay engine and observability layer.

── more in #ai-agents 4 stories · sorted by recency
── more on @burr 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/time-travel-debuggin…] indexed:0 read:6min 2026-10-06 · —