Show HN: Pawl – Hooks that check what a coding agent's script would delete Developer ulukaya released Pawl, an open-source set of deterministic hook checks for coding agents that runs as one plugin across Claude Code, OpenAI Codex and Antigravity, adding only a 144-token skill description to the prompt. Pawl's thirteen gates inspect tool calls without executing them, denying destructive commands such as a script that runs rm -rf ~/, git reset --hard, edits that change nothing, and sends naming ~/.deploy/, while auto-approving provably read-only commands like git log --oneline -5. A replay command sends every tool call from a user's past 30 days of Claude Code sessions through the same gates to tally what would have been blocked or approved without a prompt. Deterministic gates for coding agents, as one plugin for Claude Code , OpenAI Codex and Antigravity . Agents make the same mistakes over and over, and telling them not to in the prompt stops working after a page. pawl is a set of small checks that run as code, not as instructions, and a report command that tallies what they blocked. Each check watches for one mistake and refuses it, cleans it up, or for provably read-only commands waves it through without a prompt. Plain Python standard library, no model calls, nothing added to the prompt but a 144-token skill description. A pawl is the small part in a ratchet that lets the wheel move forward and stops it from slipping back. Every check here works the same way: the current state is the floor. No install, no harness, nothing written outside a scratch directory: git clone https://github.com/ulukaya/pawl && python3 pawl/hooks/pawl.py demo Recorded through the real Claude Code hook adapter using fixture tool calls. The commands are inspected, never executed; displayed reasons are excerpts. This demonstrates hook decisions rather than a live agent session. The command above sends thirteen calls through the real dispatcher, written as Antigravity, Claude Code and Codex send it, and prints what each harness is told: call gate antigravity claude code codex ------------------------------------------------------------------------------- make build - allow silent silent git log --oneline -5 readonly auto approve allow silent git reset --hard git force ask ask deny script that runs rm -rf ~/ blast deny deny deny curl install.sh | sh pin force ask ask deny tail -f server.log poll force ask ask deny same pytest run, 3rd time loop force ask ask deny edit that changes nothing noop deny deny deny write with a U+200B zero-width allow +rewrite +rewrite allow +rewrite send naming ~/.deploy/ egress deny deny deny read another session fence force ask ask deny read Chrome's cookie key creds force ask ask deny stop with tail -f running idle n/a block n/a silent leaves the harness's own prompt in place; allow and auto approve skip it. A Codex PreToolUse hook can neither ask nor approve, so there an ask is a deny that tells the agent how to proceed. --verbose adds every reason, --json every raw answer. replay sends every tool call from your past Claude Code sessions through the same gates. It runs nothing and keeps no state: python3 pawl/hooks/pawl.py replay --days 30 --show It lists each call pawl would have asked about or refused, then a tally per gate, including how many read-only commands it would have approved without a prompt. Run it before installing to see what would change, and after changing a gate to find its false alarms on real work. Grouped by what a miss costs. The first table is work or data an agent cannot take back; those gates are the reason pawl exists. Can't be undone | The mistake | What pawl does | Gate | |---|---|---| | Deletes home, root, a drive or ~/Documents, directly or from a script, trap , npm run , Makefile, python -c or container it runs | Refuses; asks before anything else outside the workspace | blast | | Runs git reset --hard , git clean -fdx , git commit --no-verify , git push --force or git branch -D and loses work | Asks the human first | git | | Pastes internal paths, tokens or hostnames into a message | Blocks the send | send egress | | Reads browser cookies, saved passwords or the keychain security find-generic-password , import browser cookie3 | Asks the human first | creds | | Reads another conversation's private files | Asks the human first | fence | | Installs from a branch URL pip install .../archive/main.zip , runs npm install -g tool with no version, or pipes curl into sh | Asks the human to pin it or approve it | pin | | Floods a chat room or inbox | Caps sends per channel per day; refuses a send loop it cannot count | send budget | Prompts it removes | The mistake | What pawl does | Gate | |---|---|---| | Stalls on a permission prompt for ls or git log | Approves commands that provably only read | readonly | Time and tokens it saves | The mistake | What pawl does | Gate | |---|---|---| | Runs while true; do sleep , tail -f or sleep 3600 and hangs | Asks the human first | poll | | Calls the same tool with the same arguments in a loop | Asks before the third identical call | loop | | Ends its turn with a tail -f still running in the background | Blocks the stop once and names the task | idle | | Changes a configured Git project without passing its checks | Reminds once at Stop; receipts match the tested contents | verify | | Re-reads its own transcript after every context truncation | Refuses past a per-turn limit | reread | | Sends an edit whose replacement equals its target, or rewrites a file with the bytes it already holds | Refuses it and sends the agent back to read | noop | | Writes invisible zero-width characters into a file | Strips them so the write lands clean | zero-width | | Writes a message that reads like a bot | Blocks it above a score threshold | send prose | CLIs for pre-commit, CI and cron | The mistake | What pawl does | Gate | |---|---|---| | Claims a bug is fixed without proving it | Requires the test to fail before the fix | repro fence.py | | Ships a test that still passes with the function stubbed out | Stubs each function in a copy and fails when the tests survive | bite check.py | | Lets failing-test or lint counts creep up | Keeps a baseline that can only go down | ratchet.py | | Grows always-on prompt files until they cost more than they help | Caps their token size | prompt budget.py | | Keeps retrying a cron job that fails every night | Pauses it after repeated failures | breaker.py | The first three tables are hook gates that fire on their own; the last holds CLIs. Every piece also runs on its own: see pieces/