# How I Built a Time-Travel Debugger for AI Agents

> Source: <https://dev.to/ujwal_bagalkoti/how-i-built-a-time-travel-debugger-for-ai-agents-b0h>
> Published: 2026-10-02 03:50:29+00:00

What if debugging an AI agent worked more like debugging normal code?

Pause execution.

Inspect what happened.

Go back to an earlier state.

Change something.

Then continue from there.

That idea led me to build an open-source **time-travel debugger for AI agents**.

GitHub: [https://github.com/UjwalBagalkoti/ai-time-travel-debugger](https://github.com/UjwalBagalkoti/ai-time-travel-debugger)

AI agents are becoming more capable, but debugging them is still surprisingly difficult.

A typical agent might execute something like:

```
User request
    ↓
LLM
    ↓
Search / Tool
    ↓
LLM
    ↓
Database / Tool
    ↓
LLM
    ↓
External API
    ↓
Final response
```

Now imagine the final answer is wrong.

The actual mistake may have happened several steps earlier.

Maybe the model selected the wrong tool.

Maybe a tool returned an unexpected result.

Maybe the agent state changed unexpectedly.

Maybe the model made a bad decision because of information introduced earlier in the execution.

With traditional logging, you can inspect what happened.

But if you want to experiment with what would have happened after changing an earlier decision, things become much harder.

You often have to run the agent again from the beginning.

That can mean:

I wanted a different approach.

Instead of thinking about an agent run as a stream of logs, I wanted to treat it as a historical execution that could be inspected.

The workflow becomes:

```
Record
   ↓
Inspect
   ↓
Rewind
   ↓
Modify
   ↓
Replay
   ↓
Branch
```

The important part is that the original execution remains available.

You can use it as the starting point for another experiment.

That's where the idea of **time-travel debugging** comes from.

The current implementation is built around a relatively simple pipeline:

```
Python AgentTracer
        ↓
JSON execution trace
        ↓
FastAPI replay API
        ↓
PostgreSQL / SQLite
        ↓
Next.js + React Flow
```

The project has four main parts.

The SDK records agent execution steps.

The backend receives, stores, retrieves, and replays traces.

PostgreSQL is used in the deployed environment, while SQLite can be used for local development.

The frontend visualizes the execution as an interactive graph and timeline.

The goal was to keep each part relatively independent so that the debugger could eventually work with different agent frameworks.

Once an execution is recorded, the frontend turns the trace into a visual execution graph.

Instead of reading a large stream of logs, you can see the individual execution steps and select them for inspection.

For example:

```
LLM
 ↓
lookup_order
 ↓
LLM
 ↓
issue_refund
```

Each step can be inspected through the debugger.

The timeline scrubber also makes it possible to move through the recorded execution history.

*Caption: The AI Time-Travel Debugger showing an agent execution graph, timeline scrubber, step inspector, and execution metrics.*

*Alt text: AI agent time-travel debugger displaying a four-step execution graph with an LLM call, tool execution, second LLM call, and final tool execution.*

This gives me a much clearer view of the execution than a traditional log file.

I can see the trajectory, select an individual step, and inspect the recorded information associated with it.

The first requirement was capturing enough information about an execution to make it useful later.

The Python tracer records information such as:

For example, a recorded tool execution can contain information like:

```
{
  "tool_name": "lookup_order",
  "arguments": {
    "order_id": "992"
  }
}
```

and its recorded result:

```
{
  "result": {
    "status": "delivered",
    "amount": 120,
    "refundable": true
  }
}
```

The important thing isn't simply collecting more logs.

The goal is to preserve enough execution history to investigate a failure.

One of the useful parts of the debugger is being able to select an individual step.

For example, selecting the `lookup_order` step shows the recorded tool arguments and the output that was produced during the original execution.

*Caption: Inspecting a recorded `lookup_order` tool execution and its recorded output.*

*Alt text: Debugger inspector showing the lookup_order tool, order ID 992, and the recorded result showing the order as delivered and refundable.*

This is important for replay because the debugger has a historical record of what the tool returned.

Instead of treating the entire execution as one opaque operation, each step can be examined separately.

This was one of the most important parts of the project.

Consider a tool call:

```
lookup_order(order_id="992")
```

During the original execution, the tool returned:

```
{
  "status": "delivered",
  "amount": 120,
  "refundable": true
}
```

If we replay the same execution and encounter the same tool with the same arguments, we don't necessarily need to execute the external tool again.

Instead, the debugger can use the recorded result.

Conceptually:

```
Original execution

lookup_order("992")
        ↓
external tool
        ↓
record result
```

Later:

```
Replay

lookup_order("992")
        ↓
recorded result
```

This is useful for both speed and safety.

It also means that debugging doesn't automatically require another network request.

The tool result is only one part of an agent execution.

The next step may be an LLM decision based on that information.

In the example execution, the next LLM step contains the instruction:

```
Authorize refund based on order details.
```

The debugger allows this intermediate step and its recorded output to be inspected.

*Caption: Inspecting the intermediate LLM step that determines whether the order is eligible for a refund.*

*Alt text: AI agent debugger showing step 3 as an LLM call with the prompt "Authorize refund based on order details" and its recorded output.*

This is where the historical execution becomes useful for debugging.

Instead of only seeing the final result, I can inspect the intermediate point where the agent made its decision.

At first, replay sounds simple.

Just execute everything again.

But that's exactly where things can go wrong.

Imagine an agent has tools such as:

```
send_email()
charge_card()
delete_database_record()
issue_refund()
```

If a debugger blindly re-executed every tool during replay, debugging could accidentally cause real-world side effects.

A debugging tool should not surprise you by sending an email or issuing a refund.

So I designed the replay system around a **safe replay boundary**.

Recorded tool calls can be replayed when their tool name and arguments match the recorded execution.

Unknown external calls aren't silently executed.

Instead, replay can stop at the boundary.

The principle is:

**Replay what is known. Don't blindly execute what isn't.**

This is still an area with many problems to solve, but I think making replay safe by default is an important foundation.

The next idea was branching.

Suppose an agent reaches this point:

```
1 → 2 → 3 → 4 → 5 → 6
```

At step 4, the agent makes a decision that leads to the wrong outcome.

With a traditional workflow, you might restart everything.

With a time-travel debugger, the goal is to create another trajectory:

```
             ┌→ 5 → 6
1 → 2 → 3 → 4
             └→ 4' → 5' → 6'
```

The original execution remains untouched.

The new execution becomes a separate branch.

The debugger exposes this through the **Fork & Replay Branch** workflow.

The historical execution effectively becomes a starting point for another experiment.

Logs are extremely useful.

But logs generally answer:

**What happened?**

A replay-oriented debugger aims to help answer another question:

**What could happen if I change something that happened earlier?**

That's the fundamental difference I'm exploring.

A normal log might tell you:

```
Step 1 completed
Step 2 completed
Step 3 failed
```

A time-travel workflow tries to give you:

```
Step 1
  ↓
Step 2
  ↓
rewind
  ↓
modify
  ↓
replay
  ↓
new execution branch
```

For increasingly complex agent systems, that distinction could become important.

Imagine an agent responsible for handling an order.

The execution is:

```
User request
 ↓
LLM
 ↓
lookup_order
 ↓
LLM
 ↓
issue_refund
 ↓
Final response
```

In my example trace, the agent looks up order `#992`.

The tool returns:

```
Status: delivered
Amount: $120
Refundable: true
```

The next LLM step evaluates the order details, and the final tool execution records the refund action.

The debugger lets each of these steps be inspected independently.

The goal is not just to see the final answer.

The goal is to understand the path that produced it.

The project currently uses:

The repository also contains a Python tracing SDK, replay engine tests, production end-to-end tests, and an example trace.

This project is still an early implementation.

There are many hard problems remaining.

LLMs aren't simple deterministic functions.

Reproducing an exact model trajectory can be difficult depending on the model, provider, configuration, tools, and surrounding state.

Some tools depend on external state.

```
current_weather()
stock_price()
web_search()
```

Their results can change between executions.

Caching their outputs helps replay, but it doesn't solve every reproducibility problem.

A real agent can modify the outside world.

That creates a fundamental question:

How should a debugger safely replay an operation that changes something outside the agent?

Agent state can become large.

Recording every state snapshot may eventually become expensive.

A production system needs strategies for:

Real agents may execute multiple operations concurrently.

A simple linear timeline isn't enough to represent every possible execution.

LLM responses and tool outputs can also be streamed.

Capturing and replaying those partial events introduces another layer of complexity.

Different agent frameworks have different execution models.

A useful debugger eventually needs to integrate naturally with the ecosystems developers already use.

The current implementation is a foundation.

The larger direction I'm interested in is making agent execution something developers can inspect, reproduce, and experiment with.

That could eventually mean deeper integrations with:

There are still many architectural questions I don't have answers to.

And that's part of why I open-sourced it.

The project is available on GitHub:

The repository contains the backend, frontend, SDK, replay engine, tests, Docker configuration, and example trace.

The deployed application is also available to experiment with:

[https://ai-time-travel-debugger.onrender.com](https://ai-time-travel-debugger.onrender.com)

I'm especially interested in feedback from developers building:

The question I'm most interested in is:

**What information must an agent runtime capture so that a failed execution can actually be reproduced and investigated?**

If you're building agents, I'd love to hear how you're currently debugging failures.

What do you record?

What do you wish you had recorded?

And where does replay break down in your systems?

AI agents are starting to look less like simple functions and more like small distributed programs.

They make decisions.

They call tools.

They maintain state.

They interact with external systems.

And when something goes wrong, "just run it again" isn't always enough.

That's the problem I'm exploring with this project.

**Record once.**

**Inspect the past.**

**Rewind the execution.**

**Experiment with another path.**

**Replay safely.**
