Show HN: Runtape – counterfactual debugging and regression tests for AI agents RehanMohammed985 released Runtape, an open-source Python tool that performs counterfactual debugging and regression testing on AI agent runs, installable via pip for Python 3.10+ and compatible with the OpenAI and Anthropic SDKs, LangChain, LangGraph, and OpenAI-compatible local servers. In a documented example, Runtape traced an email assistant's forwarding of an invoice to a single sentence in an HTML comment inside a vendor email, showing the agent forwarded in 10 of 10 reruns with the sentence present and 0 of 10 without it (p = 5e-6), then wrote a pytest regression test for the fix that held. A separate real case on llama3.2 (3B) found a refund agent paid order B-2290 $64, an amount from a different customer's earlier order, in 9 of 40 reruns, dropping to 0 of 40 once the earlier order lookup was removed (p = 0.001). Counterfactual debugging and regression tests for AI agents. Give runtape a bad agent run. It finds the part of the context that caused the bad decision, checks candidate fixes against the exact context that failed, and writes a regression test so it stays fixed. runtape why last tool:forward email what caused it runtape fix last tool:forward email which fixes hold, measured runtape fix last tool:forward email --write-test tests/test inbox.py An email assistant forwards an invoice to an outside address. runtape why traces the call to one sentence in an HTML comment inside a vendor email: with it, the agent forwards in 10 of 10 reruns; without it, in 0 of 10 p = 5e-6 . runtape fix then tries system prompt rules and fixing the source, reruns the decision with each, and writes a pytest file for the fix that holds. Tracing tools such as LangSmith and Langfuse show what the agent saw. Attribution methods such as ContextCite score context for a single model response. runtape works on your agent's own recorded runs, on your machine, and is meant for investigating a specific failure and keeping it fixed. pip install runtape Python 3.10+. Works with the OpenAI and Anthropic SDKs, LangChain and LangGraph, OpenAI-compatible local servers Ollama, LM Studio, vLLM , and custom agent loops. git clone https://github.com/RehanMohammed985/runtape cd runtape pip install . openai anthropic python examples/inbox agent.py runtape why last tool:forward email --model-fn examples/inbox agent.py:simulated model runtape fix last tool:forward email --model-fn examples/inbox agent.py:simulated model --write-test tests/test inbox.py pytest tests/test inbox.py | example | failure | |---|---| | inbox agent.py | an email assistant forwards an invoice because of an instruction hidden in an email | | refund bot.py | a support agent refunds $2,400 after reading a stale forum post in search results | | ops agent.py | an operations agent drops a shared staging database, following an old runbook line | By default the examples run offline with a rule-based stand-in model --model-fn . To run them on a real model, add --local MODEL Ollama, free , --openai MODEL or --anthropic MODEL . Real models don't fail every time, so examples/hunt.py runs an example until it fails, reports the tokens used, and prints the why command: python examples/hunt.py ops --local llama3.1:8b --tries 5 A real case on llama3.2 3B : the refund agent paid order B-2290 $64, the amount from a different customer's order earlier in the conversation. On the recorded context it did this in 9 of 40 reruns; with the earlier order lookup removed, in 0 of 40 p = 0.001 . python import runtape from openai import OpenAI rec = runtape.record name="support-bot" client = rec.wrap OpenAI every model call is recorded @rec.tool arguments, results, errors, latency def lookup order order id: str : ... Traces go to ./traces/ , one JSONL file per run. See docs/usage.md https://github.com/RehanMohammed985/runtape/blob/main/docs/usage.md for Anthropic, LangChain, streaming and custom loops. runtape why