{"slug": "how-i-built-a-time-travel-debugger-for-ai-agents", "title": "How I Built a Time-Travel Debugger for AI Agents", "summary": "A developer built an open-source time-travel debugger for AI agents that records execution traces, lets users rewind to earlier states, modify decisions, and replay or branch from any step. The project pairs a Python AgentTracer SDK with a FastAPI replay API, PostgreSQL or SQLite storage, and a Next.js/React Flow frontend that renders agent runs as interactive execution graphs and timelines. The goal is to make agent debugging work more like conventional code debugging instead of re-running an agent from the beginning.", "body_md": "What if debugging an AI agent worked more like debugging normal code?\n\nPause execution.\n\nInspect what happened.\n\nGo back to an earlier state.\n\nChange something.\n\nThen continue from there.\n\nThat idea led me to build an open-source **time-travel debugger for AI agents**.\n\nGitHub: [https://github.com/UjwalBagalkoti/ai-time-travel-debugger](https://github.com/UjwalBagalkoti/ai-time-travel-debugger)\n\nAI agents are becoming more capable, but debugging them is still surprisingly difficult.\n\nA typical agent might execute something like:\n\n```\nUser request\n    ↓\nLLM\n    ↓\nSearch / Tool\n    ↓\nLLM\n    ↓\nDatabase / Tool\n    ↓\nLLM\n    ↓\nExternal API\n    ↓\nFinal response\n```\n\nNow imagine the final answer is wrong.\n\nThe actual mistake may have happened several steps earlier.\n\nMaybe the model selected the wrong tool.\n\nMaybe a tool returned an unexpected result.\n\nMaybe the agent state changed unexpectedly.\n\nMaybe the model made a bad decision because of information introduced earlier in the execution.\n\nWith traditional logging, you can inspect what happened.\n\nBut if you want to experiment with what would have happened after changing an earlier decision, things become much harder.\n\nYou often have to run the agent again from the beginning.\n\nThat can mean:\n\nI wanted a different approach.\n\nInstead of thinking about an agent run as a stream of logs, I wanted to treat it as a historical execution that could be inspected.\n\nThe workflow becomes:\n\n```\nRecord\n   ↓\nInspect\n   ↓\nRewind\n   ↓\nModify\n   ↓\nReplay\n   ↓\nBranch\n```\n\nThe important part is that the original execution remains available.\n\nYou can use it as the starting point for another experiment.\n\nThat's where the idea of **time-travel debugging** comes from.\n\nThe current implementation is built around a relatively simple pipeline:\n\n```\nPython AgentTracer\n        ↓\nJSON execution trace\n        ↓\nFastAPI replay API\n        ↓\nPostgreSQL / SQLite\n        ↓\nNext.js + React Flow\n```\n\nThe project has four main parts.\n\nThe SDK records agent execution steps.\n\nThe backend receives, stores, retrieves, and replays traces.\n\nPostgreSQL is used in the deployed environment, while SQLite can be used for local development.\n\nThe frontend visualizes the execution as an interactive graph and timeline.\n\nThe goal was to keep each part relatively independent so that the debugger could eventually work with different agent frameworks.\n\nOnce an execution is recorded, the frontend turns the trace into a visual execution graph.\n\nInstead of reading a large stream of logs, you can see the individual execution steps and select them for inspection.\n\nFor example:\n\n```\nLLM\n ↓\nlookup_order\n ↓\nLLM\n ↓\nissue_refund\n```\n\nEach step can be inspected through the debugger.\n\nThe timeline scrubber also makes it possible to move through the recorded execution history.\n\n*Caption: The AI Time-Travel Debugger showing an agent execution graph, timeline scrubber, step inspector, and execution metrics.*\n\n*Alt text: AI agent time-travel debugger displaying a four-step execution graph with an LLM call, tool execution, second LLM call, and final tool execution.*\n\nThis gives me a much clearer view of the execution than a traditional log file.\n\nI can see the trajectory, select an individual step, and inspect the recorded information associated with it.\n\nThe first requirement was capturing enough information about an execution to make it useful later.\n\nThe Python tracer records information such as:\n\nFor example, a recorded tool execution can contain information like:\n\n```\n{\n  \"tool_name\": \"lookup_order\",\n  \"arguments\": {\n    \"order_id\": \"992\"\n  }\n}\n```\n\nand its recorded result:\n\n```\n{\n  \"result\": {\n    \"status\": \"delivered\",\n    \"amount\": 120,\n    \"refundable\": true\n  }\n}\n```\n\nThe important thing isn't simply collecting more logs.\n\nThe goal is to preserve enough execution history to investigate a failure.\n\nOne of the useful parts of the debugger is being able to select an individual step.\n\nFor example, selecting the `lookup_order` step shows the recorded tool arguments and the output that was produced during the original execution.\n\n*Caption: Inspecting a recorded `lookup_order` tool execution and its recorded output.*\n\n*Alt text: Debugger inspector showing the lookup_order tool, order ID 992, and the recorded result showing the order as delivered and refundable.*\n\nThis is important for replay because the debugger has a historical record of what the tool returned.\n\nInstead of treating the entire execution as one opaque operation, each step can be examined separately.\n\nThis was one of the most important parts of the project.\n\nConsider a tool call:\n\n```\nlookup_order(order_id=\"992\")\n```\n\nDuring the original execution, the tool returned:\n\n```\n{\n  \"status\": \"delivered\",\n  \"amount\": 120,\n  \"refundable\": true\n}\n```\n\nIf we replay the same execution and encounter the same tool with the same arguments, we don't necessarily need to execute the external tool again.\n\nInstead, the debugger can use the recorded result.\n\nConceptually:\n\n```\nOriginal execution\n\nlookup_order(\"992\")\n        ↓\nexternal tool\n        ↓\nrecord result\n```\n\nLater:\n\n```\nReplay\n\nlookup_order(\"992\")\n        ↓\nrecorded result\n```\n\nThis is useful for both speed and safety.\n\nIt also means that debugging doesn't automatically require another network request.\n\nThe tool result is only one part of an agent execution.\n\nThe next step may be an LLM decision based on that information.\n\nIn the example execution, the next LLM step contains the instruction:\n\n```\nAuthorize refund based on order details.\n```\n\nThe debugger allows this intermediate step and its recorded output to be inspected.\n\n*Caption: Inspecting the intermediate LLM step that determines whether the order is eligible for a refund.*\n\n*Alt text: AI agent debugger showing step 3 as an LLM call with the prompt \"Authorize refund based on order details\" and its recorded output.*\n\nThis is where the historical execution becomes useful for debugging.\n\nInstead of only seeing the final result, I can inspect the intermediate point where the agent made its decision.\n\nAt first, replay sounds simple.\n\nJust execute everything again.\n\nBut that's exactly where things can go wrong.\n\nImagine an agent has tools such as:\n\n```\nsend_email()\ncharge_card()\ndelete_database_record()\nissue_refund()\n```\n\nIf a debugger blindly re-executed every tool during replay, debugging could accidentally cause real-world side effects.\n\nA debugging tool should not surprise you by sending an email or issuing a refund.\n\nSo I designed the replay system around a **safe replay boundary**.\n\nRecorded tool calls can be replayed when their tool name and arguments match the recorded execution.\n\nUnknown external calls aren't silently executed.\n\nInstead, replay can stop at the boundary.\n\nThe principle is:\n\n**Replay what is known. Don't blindly execute what isn't.**\n\nThis is still an area with many problems to solve, but I think making replay safe by default is an important foundation.\n\nThe next idea was branching.\n\nSuppose an agent reaches this point:\n\n```\n1 → 2 → 3 → 4 → 5 → 6\n```\n\nAt step 4, the agent makes a decision that leads to the wrong outcome.\n\nWith a traditional workflow, you might restart everything.\n\nWith a time-travel debugger, the goal is to create another trajectory:\n\n```\n             ┌→ 5 → 6\n1 → 2 → 3 → 4\n             └→ 4' → 5' → 6'\n```\n\nThe original execution remains untouched.\n\nThe new execution becomes a separate branch.\n\nThe debugger exposes this through the **Fork & Replay Branch** workflow.\n\nThe historical execution effectively becomes a starting point for another experiment.\n\nLogs are extremely useful.\n\nBut logs generally answer:\n\n**What happened?**\n\nA replay-oriented debugger aims to help answer another question:\n\n**What could happen if I change something that happened earlier?**\n\nThat's the fundamental difference I'm exploring.\n\nA normal log might tell you:\n\n```\nStep 1 completed\nStep 2 completed\nStep 3 failed\n```\n\nA time-travel workflow tries to give you:\n\n```\nStep 1\n  ↓\nStep 2\n  ↓\nrewind\n  ↓\nmodify\n  ↓\nreplay\n  ↓\nnew execution branch\n```\n\nFor increasingly complex agent systems, that distinction could become important.\n\nImagine an agent responsible for handling an order.\n\nThe execution is:\n\n```\nUser request\n ↓\nLLM\n ↓\nlookup_order\n ↓\nLLM\n ↓\nissue_refund\n ↓\nFinal response\n```\n\nIn my example trace, the agent looks up order `#992`.\n\nThe tool returns:\n\n```\nStatus: delivered\nAmount: $120\nRefundable: true\n```\n\nThe next LLM step evaluates the order details, and the final tool execution records the refund action.\n\nThe debugger lets each of these steps be inspected independently.\n\nThe goal is not just to see the final answer.\n\nThe goal is to understand the path that produced it.\n\nThe project currently uses:\n\nThe repository also contains a Python tracing SDK, replay engine tests, production end-to-end tests, and an example trace.\n\nThis project is still an early implementation.\n\nThere are many hard problems remaining.\n\nLLMs aren't simple deterministic functions.\n\nReproducing an exact model trajectory can be difficult depending on the model, provider, configuration, tools, and surrounding state.\n\nSome tools depend on external state.\n\n```\ncurrent_weather()\nstock_price()\nweb_search()\n```\n\nTheir results can change between executions.\n\nCaching their outputs helps replay, but it doesn't solve every reproducibility problem.\n\nA real agent can modify the outside world.\n\nThat creates a fundamental question:\n\nHow should a debugger safely replay an operation that changes something outside the agent?\n\nAgent state can become large.\n\nRecording every state snapshot may eventually become expensive.\n\nA production system needs strategies for:\n\nReal agents may execute multiple operations concurrently.\n\nA simple linear timeline isn't enough to represent every possible execution.\n\nLLM responses and tool outputs can also be streamed.\n\nCapturing and replaying those partial events introduces another layer of complexity.\n\nDifferent agent frameworks have different execution models.\n\nA useful debugger eventually needs to integrate naturally with the ecosystems developers already use.\n\nThe current implementation is a foundation.\n\nThe larger direction I'm interested in is making agent execution something developers can inspect, reproduce, and experiment with.\n\nThat could eventually mean deeper integrations with:\n\nThere are still many architectural questions I don't have answers to.\n\nAnd that's part of why I open-sourced it.\n\nThe project is available on GitHub:\n\nThe repository contains the backend, frontend, SDK, replay engine, tests, Docker configuration, and example trace.\n\nThe deployed application is also available to experiment with:\n\n[https://ai-time-travel-debugger.onrender.com](https://ai-time-travel-debugger.onrender.com)\n\nI'm especially interested in feedback from developers building:\n\nThe question I'm most interested in is:\n\n**What information must an agent runtime capture so that a failed execution can actually be reproduced and investigated?**\n\nIf you're building agents, I'd love to hear how you're currently debugging failures.\n\nWhat do you record?\n\nWhat do you wish you had recorded?\n\nAnd where does replay break down in your systems?\n\nAI agents are starting to look less like simple functions and more like small distributed programs.\n\nThey make decisions.\n\nThey call tools.\n\nThey maintain state.\n\nThey interact with external systems.\n\nAnd when something goes wrong, \"just run it again\" isn't always enough.\n\nThat's the problem I'm exploring with this project.\n\n**Record once.**\n\n**Inspect the past.**\n\n**Rewind the execution.**\n\n**Experiment with another path.**\n\n**Replay safely.**", "url": "https://wpnews.pro/news/how-i-built-a-time-travel-debugger-for-ai-agents", "canonical_source": "https://dev.to/ujwal_bagalkoti/how-i-built-a-time-travel-debugger-for-ai-agents-b0h", "published_at": "2026-10-02 03:50:29+00:00", "updated_at": "2026-10-02 04:14:38.298130+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "mlops", "artificial-intelligence"], "entities": ["UjwalBagalkoti", "Python", "FastAPI", "PostgreSQL", "SQLite", "Next.js", "React Flow"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-i-built-a-time-travel-debugger-for-ai-agents", "markdown": "https://wpnews.pro/news/how-i-built-a-time-travel-debugger-for-ai-agents.md", "text": "https://wpnews.pro/news/how-i-built-a-time-travel-debugger-for-ai-agents.txt", "jsonld": "https://wpnews.pro/news/how-i-built-a-time-travel-debugger-for-ai-agents.jsonld"}}