I replayed two months of my Claude Code sessions. $1,675 at API prices, and 117 rm -rf. A developer built Paveo, a policy checkpoint that intercepts and blocks unsafe tool calls before an AI agent executes them, and replayed two months of their own Claude Code session logs to test it. Across 59 days, 106 sessions, 8,113 model calls and 8,176 tool calls, the usage would have cost $1,674.75 at API list prices, with a single session reaching $452.92, and the starter policy would have refused 670 tool calls including 117 recursive force deletes and 547 edits to the agent's own settings or the guard's files. The developer notes many refusals were false alarms but kept the rule guarding the guard, and Paveo's repository is going public this week, running in-process with no network sockets and supporting Claude Code, Codex, Cursor or custom agents. I build Paveo, a checkpoint that sits between an AI agent and everything it calls and refuses the calls it shouldn't make, before they run. Before I asked anyone else to trust that, I ran it against the agent I use most: my own Claude Code, over the two months I spent building Paveo. I'm on a subscription, so none of this was billed. That matters for everything below: these are the numbers my usage would have cost on the API, and the commands my agent would have been stopped from running. Claude Code keeps every session on your disk, one .jsonl file per session under ~/.claude/projects . Each assistant message records its token usage, and each tool call records its name and arguments. paveo replay reads those files, prices the model calls from a dated price table, and puts every tool call through a policy. It stores nothing and prints only counts and rule names. 59 days. 106 sessions. 8,113 model calls. 8,176 tool calls. | What | Amount | |---|---| | Total at API list prices | $1,674.75 | | Costliest day | $136.93 | | Costliest single session | $452.92 | | Opus 5 | $1,300.47 over 5,248 calls | | Opus 5.5 | $374.28 over 2,861 calls | Almost every call in the files omits which region served it, and US-only inference costs more on some models, so those calls are priced at the US rate. The total can read up to 10% high. It never reads low, which is the direction a budget should err in. The number that stayed with me is the single session: $452.92. On a subscription it was an ordinary long afternoon. On the API, as an unattended agent, it's the incident report. The starter policy for Claude Code is short: a handful of patterns for commands that destroy work, and the paths to the agent's own settings and the guard's own files. Replayed over 8,176 tool calls, it would have refused 670: | Refused | Count | |---|---| | rm -rf recursive force delete | 117 | | git clean -f | 3 | | git reset --hard | 2 | | git push --force | 1 | | edits to .claude settings or the guard's own files | 547 | Two honest notes on those numbers. Not every rm -rf was a mistake. Plenty were the agent cleaning up a build folder it had just made. The guard would have made it find another way, or ask. I'm fine with that trade: a false refusal costs a retry, a missed one costs the work. Most of the 547 are false alarms, and I kept the rule anyway. I was building a tool whose files live in .paveo , so the agent touched them constantly. But "the agent edits its own settings or its own guard" is the one thing I least want to allow, because it's how every other limit gets switched off. A noisy rule that guards the guard stays. Paveo's repository goes public this week. It runs in your own process, opens no network sockets, and the same policy file guards Claude Code, Codex, Cursor or an agent you build yourself. If you've priced your own sessions, I'd like to know: what was your costliest one?