# Show HN: Pawl – Hooks that check what a coding agent's script would delete

> Source: <https://github.com/ulukaya/pawl>
> Published: 2026-10-06 15:09:07+00:00

Deterministic gates for coding agents, as one plugin for **Claude Code**,
**OpenAI Codex** and **Antigravity**.

Agents make the same mistakes over and over, and telling them not to in the
prompt stops working after a page. `pawl` is a set of small checks that run as
code, not as instructions, and a report command that tallies what they blocked.
Each check watches for one mistake and refuses it, cleans it up, or (for
provably read-only commands) waves it through without a prompt. Plain Python
standard library, no model calls, nothing added to the prompt but a 144-token
skill description.

A pawl is the small part in a ratchet that lets the wheel move forward and stops it from slipping back. Every check here works the same way: the current state is the floor.

No install, no harness, nothing written outside a scratch directory:

```
git clone https://github.com/ulukaya/pawl && python3 pawl/hooks/pawl.py demo
```

Recorded through the real Claude Code hook adapter using fixture tool calls. The commands are inspected, never executed; displayed reasons are excerpts. This demonstrates hook decisions rather than a live agent session.

The command above sends thirteen calls through the real dispatcher, written as Antigravity, Claude Code and Codex send it, and prints what each harness is told:

```
call                        gate        antigravity      claude code      codex
-------------------------------------------------------------------------------
make build                  -           allow            silent           silent
git log --oneline -5        readonly    auto_approve     allow            silent
git reset --hard            git         force_ask        ask              deny
script that runs rm -rf ~/  blast       deny             deny             deny
curl install.sh | sh        pin         force_ask        ask              deny
tail -f server.log          poll        force_ask        ask              deny
same pytest run, 3rd time   loop        force_ask        ask              deny
edit that changes nothing   noop        deny             deny             deny
write with a U+200B         zero-width  allow +rewrite   +rewrite         allow +rewrite
send naming ~/.deploy/      egress      deny             deny             deny
read another session        fence       force_ask        ask              deny
read Chrome's cookie key    creds       force_ask        ask              deny
stop with tail -f running   idle        n/a              block            n/a
```

`silent` leaves the harness's own prompt in place; `allow` and
`auto_approve` skip it. A Codex PreToolUse hook can neither ask nor approve,
so there an ask is a deny that tells the agent how to proceed. `--verbose` adds every
reason, `--json` every raw answer.

`replay` sends every tool call from your past Claude Code sessions through
the same gates. It runs nothing and keeps no state:

```
python3 pawl/hooks/pawl.py replay --days 30 --show
```

It lists each call pawl would have asked about or refused, then a tally per gate, including how many read-only commands it would have approved without a prompt. Run it before installing to see what would change, and after changing a gate to find its false alarms on real work.

Grouped by what a miss costs. The first table is work or data an agent cannot take back; those gates are the reason pawl exists.

**Can't be undone**

| The mistake | What pawl does | Gate | 
|---|---|---|
| Deletes home, root, a drive or ~/Documents, directly or from a script, `trap` ,`npm run` , Makefile,`python -c` or container it runs | Refuses; asks before anything else outside the workspace | `blast` | 
| Runs `git reset --hard` ,`git clean -fdx` ,`git commit --no-verify` ,`git push --force` or`git branch -D` and loses work | Asks the human first | `git` | 
| Pastes internal paths, tokens or hostnames into a message | Blocks the send | `send` (egress) | 
| Reads browser cookies, saved passwords or the keychain ( `security find-generic-password` ,`import browser_cookie3` ) | Asks the human first | `creds` | 
| Reads another conversation's private files | Asks the human first | `fence` | 
| Installs from a branch URL ( `pip install .../archive/main.zip` ), runs`npm install -g tool` with no version, or pipes`curl` into`sh` | Asks the human to pin it or approve it | `pin` | 
| Floods a chat room or inbox | Caps sends per channel per day; refuses a send loop it cannot count | `send` (budget) | 

**Prompts it removes**

| The mistake | What pawl does | Gate | 
|---|---|---|
| Stalls on a permission prompt for `ls` or`git log` | Approves commands that provably only read | `readonly` | 

**Time and tokens it saves**

| The mistake | What pawl does | Gate | 
|---|---|---|
| Runs `while true; do sleep` ,`tail -f` or`sleep 3600` and hangs | Asks the human first | `poll` | 
| Calls the same tool with the same arguments in a loop | Asks before the third identical call | `loop` | 
| Ends its turn with a `tail -f` still running in the background | Blocks the stop once and names the task | `idle` | 
| Changes a configured Git project without passing its checks | Reminds once at Stop; receipts match the tested contents | `verify` | 
| Re-reads its own transcript after every context truncation | Refuses past a per-turn limit | `reread` | 
| Sends an edit whose replacement equals its target, or rewrites a file with the bytes it already holds | Refuses it and sends the agent back to read | `noop` | 
| Writes invisible zero-width characters into a file | Strips them so the write lands clean | `zero-width` | 
| Writes a message that reads like a bot | Blocks it above a score threshold | `send` (prose) | 

**CLIs for pre-commit, CI and cron**

| The mistake | What pawl does | Gate | 
|---|---|---|
| Claims a bug is fixed without proving it | Requires the test to fail before the fix | `repro_fence.py` | 
| Ships a test that still passes with the function stubbed out | Stubs each function in a copy and fails when the tests survive | `bite_check.py` | 
| Lets failing-test or lint counts creep up | Keeps a baseline that can only go down | `ratchet.py` | 
| Grows always-on prompt files until they cost more than they help | Caps their token size | `prompt_budget.py` | 
| Keeps retrying a cron job that fails every night | Pauses it after repeated failures | `breaker.py` | 

The first three tables are hook gates that fire on their own; the
last holds CLIs. Every piece also runs on its own: see
`pieces/<name>/README.md`.

| Prompt | 144 tokens: the skill's description, the only always-on text | 
| Latency | about 45 ms per Read and 65 ms per Bash call (median, Linux, Python 3.11), all gates in one process | 
| Network | none: no telemetry, no model calls ( [PRIVACY.md](https://github.com/ulukaya/pawl/blob/main/PRIVACY.md) ) | 
| Dependencies | the Python 3.11+ standard library | 

```
claude plugin marketplace add ulukaya/pawl
claude plugin install pawl@pawl
```

Or inside a session: `/plugin marketplace add ulukaya/pawl`, then
`/plugin install pawl@pawl`. Start a new session (or `/reload-plugins`), and
`claude plugin details pawl` lists `Hooks (2) PreToolUse, Stop`.

```
codex plugin marketplace add ulukaya/pawl
codex plugin add pawl@pawl
```

Codex reads `.codex-plugin/plugin.json`, which points it at
`hooks/codex.json` and the skill. Hooks need a Codex release with lifecycle
hooks.

Adding the plugin does not turn its hooks on. Codex runs a plugin's hooks
only once you have reviewed and trusted their current definitions, and skips
new or changed ones until then; see
[review and trust hooks](https://developers.openai.com/codex/hooks#review-and-trust-hooks)
in the Codex docs. To activate pawl:

1. Start `codex` . It warns at startup when hooks need review.
2. Open `/hooks` , review pawl's`PreToolUse` and`Stop` hooks (each runs`hooks/pawl.py` with`--harness codex` ) and trust both.
3. Do it again after an update that changes `hooks/codex.json` : a changed
definition is skipped until trusted again.

To check that pawl is live, open `/hooks` and confirm both hooks are trusted
and enabled. Then, in a scratch session, ask Codex to run `echo pawl-check`
three times as separate commands: the third is refused with a
`[PAWL loop]` reason. `hooks/e2e_test.py` does not prove this: it runs the
shipped hook commands directly on fixture payloads, with no Codex process,
so it passes whether or not Codex loaded or trusted the hooks.

```
git clone https://github.com/ulukaya/pawl && cd pawl
./install.sh --antigravity
```

This links the checkout to `~/.gemini/config/plugins/pawl` and adds
`{"path": "plugins/pawl"}` to `~/.gemini/config/plugins.json` next to any
plugins already listed. Restart Antigravity; the plugin inventory lists
`pawl`.

```
./install.sh                 # every harness found on this machine
./install.sh --claude        # or --antigravity, --codex
./install.sh --uninstall     # reverse every step
./install.sh --dry-run       # print what would change
```

For Claude Code and Codex the installer runs the commands above with this
checkout as the marketplace, so the plugin loads in place (Codex still needs
the `/hooks` trust step above; the installer reminds you); `--source ulukaya/pawl` tracks GitHub instead. Every step is idempotent. It then runs
the test battery, which needs `pytest` (`PAWL_PYTHON` picks the
interpreter); the plugin itself needs only Python 3.11+.

```
 harness ──stdin──▶ hooks/pawl.py pre|stop --harness H
                      │
                      ├─ harness.parse()    native payload ─▶ one canonical call
                      ├─ gates.plan()       which gates apply to this tool
                      ├─ pieces/<name>/     each gate asks its piece
                      ├─ merge              deny > ask > approve > allow
                      └─ harness.render()   answer in H's own contract ──stdout──▶
```

Each harness has its own config, all running the same dispatcher:

| Harness | Config | Command | 
|---|---|---|
| Antigravity | `hooks.json` | `python3 -B hooks/pawl.py pre --only <gate> --harness antigravity` , one group per gate | 
| Claude Code | `hooks/hooks.json` | `python3 -B "${CLAUDE_PLUGIN_ROOT}/hooks/pawl.py" pre --harness claude` , plus`stop` | 
| Codex | `hooks/codex.json` | `python3 -B "${PLUGIN_ROOT}/hooks/pawl.py" pre --harness codex` , plus`stop` | 

`hooks/harness.py`, with its tables in `hooks/harness_vocab.py`, maps each
harness's tools onto the canonical names the pieces speak (Claude Code
`Bash`, `Read`, `Write`, `Edit`, `TaskOutput`; Codex `Bash` and
`apply_patch`) and writes each answer the way that harness reads it:

| pawl decides | Antigravity | Claude Code | Codex | 
|---|---|---|---|
| no objection | `allow` | no output: the normal permission prompt still applies | no output | 
| ask the human | `force_ask` | `permissionDecision: ask` | `deny` with the reason: a PreToolUse hook cannot ask | 
| refuse | `deny` | `permissionDecision: deny` | `permissionDecision: deny` | 
| provably read-only | `auto_approve` | `permissionDecision: allow` | no output: a PreToolUse hook cannot approve | 
| rewrite the input | `overwrite` | `updatedInput` , permission unchanged | `updatedInput` | 
| keep working (Stop) | `block` | `decision: block` | `decision: block` | 

`git`, `poll`, `pin`, `creds` and `fence` read a shell command through
`hooks/shell_view.py`, which blanks heredoc bodies that never run as shell:
the file `cat > notes.md <<'EOF'` writes, a commit message in
`git commit -m "$(cat <<'EOF' ...)"`, the string literals of a
`python3 - <<'EOF'` edit script, a list of commands that a `while read`
loop or a later `python3 check.py cases.txt` only reads. Each rule names
what is known to be data, and anything else keeps the body: a file run or
copied later, a loop that runs or saves the lines it reads, and Python
whose code can start a process (`hooks/py_body.py` reads it with `ast`).
`creds` and `fence` keep the paths in Python literals.

Codex does let a separate `PermissionRequest` hook, sent only when Codex is
about to ask the user, answer allow or deny; pawl registers no such hook.

Reasons are reworded in the harness's own tool names (`Read`, not
`view_file`). The first deny ends a run, so a refused call never spends a
send-budget unit. On Codex an ask is a deny, so it ends the run as well, and
a send over budget is refused with nothing spent and no override logged.
A git or poll gate whose piece cannot load denies, as it does when it fails.
Every answer exits 0; no path prints a traceback.

| Gate | Fires on | Fails | Antigravity | Claude Code | Codex | 
|---|---|---|---|---|---|
| `fence` | every tool | open | `brain/` ,`conversations/` | `~/.claude/projects/` | `~/.codex/sessions/` | 
| `creds` | every tool | open | yes | yes | yes (ask is deny) | 
| `git` | shell | closed | yes | yes | yes (deny) | 
| `blast` | shell | open (an unfinished analysis asks) | yes | yes | yes (ask is deny) | 
| `pin` | shell | open | yes | yes | yes (ask is deny) | 
| `poll` | shell | closed | yes | yes | yes (deny) | 
| `noop` | edits | open | `replace_file_content` | `Edit` | `apply_patch` | 
| `zero-width` | writes, edits | open | yes | `Write` ,`Edit` | `apply_patch` | 
| `readonly` | shell | open | yes | yes | no (PreToolUse cannot approve) | 
| `reread` | reads, shell | open | yes | yes | yes | 
| `loop` | every tool | open | yes | yes | yes (deny) | 
| `send` | shell | egress closed | yes | yes | yes | 
| `idle` | Stop | open | `/proc` tasks | Stop payload tasks | no task list | 
| `verify` | all + Stop | open | configured Git files | configured Git files | configured Git files | 

`hooks/pawl.py gates` lists them. Details: `skills/pawl/references/<piece>.md`
and `pieces/<piece>/README.md`.

| Need | Do | 
|---|---|
| Skip gates for a session | `PAWL_DISABLE=git,poll` (any gate name, or`egress` ,`prose` ,`budget` ) | 
| See what pawl would have done in past sessions | `python3 hooks/pawl.py replay --days 30 --show` (Claude Code transcripts) | 
| See how often each send gate fires | `python3 hooks/pawl.py stats` (reads`$PAWL_DATA/gate_events.jsonl` ) | 
| Tally every gate's denials | `python3 pieces/report/report.py --days 7` | 
| One send, git or repeated call past a gate | approve the prompt; the row is logged as a human override | 
| Change send ceilings | `SEND_BUDGET_CEILINGS='{"chat_space": 4}'` | 
| Change egress rules | edit `$PAWL_DATA/egress_rules.json` (default`~/.pawl/` ) | 
| Protect specific repos | `PAWL_GIT_PROTECTED_ROOTS=/repo/a:/repo/b` (default: the call's git toplevel) | 
| Keep the prompt for read-only commands | `PAWL_READONLY_PASS_OFF=1` | 
| Refuse, not ask, on another session's files | `PAWL_CONVERSATION_FENCE_STRICT=1` | 
| Fence another credential store | `PAWL_CREDENTIAL_EXTRA_ROOTS=~/.agent-reach` | 
| Refuse, not ask, on browser logins or unpinned installs | `PAWL_CREDENTIAL_FENCE_STRICT=1` ,`PAWL_INSTALL_PIN_GUARD_STRICT=1` | 
| Force a harness format | `--harness` in the config, or`PAWL_HARNESS` | 

On Claude Code four of these are plugin settings, so nobody edits an
environment: `/config` lists pawl's rows (Approve read-only shell commands,
Repos the git guard protects, Refuse reads of other sessions, Gates to turn
off), and `claude plugin install pawl@pawl --config disable=reread` sets one
at install. A `PAWL_*` variable you export wins over the setting.

State and logs live under `PAWL_DATA` (default `~/.pawl`). Each piece's knobs
are listed in its reference page.

pawl runs on your machine and nowhere else: no network, no telemetry, no
model calls. Its logs hold counters and SHA-1 digests, never a command,
path or message. One gate loosens anything: on Claude Code and Antigravity,
`readonly` approves shell commands it can prove only read, without a prompt
(a Codex PreToolUse hook cannot approve, so there it stays silent); turn it
off in `/config` or with `PAWL_READONLY_PASS_OFF=1`.
[PRIVACY.md](https://github.com/ulukaya/pawl/blob/main/PRIVACY.md) lists every file pawl reads and writes.

| Piece | Wire point | Command | 
|---|---|---|
| prose-gate | gate `send` , CI on docs | `pieces/prose-gate/prose_gate.py --plane chat draft.md` | 
| egress-firewall | gate `send` | `pieces/egress-firewall/egress_firewall.py check < text` | 
| send-budget | gate `send` | `pieces/send-budget/send_budget.py status` | 
| blast-radius-guard | gate `blast` | `pieces/blast-radius-guard/blast_radius.py check --cwd . -- bash cleanup.sh` | 
| destructive-git-guard | gate `git` | `pieces/destructive-git-guard/destructive_git_guard.py check --cwd . git reset --hard` | 
| poll-loop-guard | gate `poll` | `pieces/poll-loop-guard/poll_loop_guard.py classify "while true; do sleep 5; done"` | 
| oscillation-breaker | gate `loop` | `pieces/oscillation-breaker/oscillation_breaker.py check <conversation-id> view_file '{"path": "a"}'` | 
| noop-edit-guard | gate `noop` | `pieces/noop-edit-guard/noop_edit_guard.py check replace_file_content '{"TargetContent": "a", "ReplacementContent": "a"}'` | 
| zero-width-sanitizer | gate `zero-width` | `pieces/zero-width-sanitizer/zero_width_sanitizer.py strip < draft.txt` | 
| readonly-pass | gate `readonly` | `pieces/readonly-pass/readonly_pass.py check git log -5` | 
| reread-guard | gate `reread` | `pieces/reread-guard/reread_guard_hook.py < payload.json` | 
| conversation-fence | gate `fence` | `pieces/conversation-fence/conversation_fence_hook.py < payload.json` | 
| credential-fence | gate `creds` | `pieces/credential-fence/credential_fence.py check security dump-keychain` | 
| install-pin-guard | gate `pin` | `pieces/install-pin-guard/install_pin_guard.py check npm install -g mcporter` | 
| idle-task-gate | gate `idle` (Stop) | `pieces/idle-task-gate/idle_task_gate.py list <conversation-id>` | 
| ratchet-baseline | pre-commit, CI | `pieces/ratchet-baseline/ratchet.py check --baseline .ratchet.json --metric failing_tests=N` | 
| circuit-breaker | cron, sidecars | `pieces/circuit-breaker/breaker.py run nightly -- ./job.sh` | 
| repro-fence | pre-commit on bug fixes | `pieces/repro-fence/repro_fence.py both --cmd "pytest tests/test_x.py" --file src/x.py` | 
| bite-check | pre-commit, CI on new tests | `pieces/bite-check/bite_check.py check --file src/x.py --cmd "pytest tests/test_x.py"` | 
| prompt-budget | pre-commit on prompt files | `pieces/prompt-budget/prompt_budget.py check --config budget.json` | 
| report | CLI, retro | `pieces/report/report.py --days 7` | 

One skill, `skills/pawl/SKILL.md`, routes an agent by denial prefix or task to
`skills/pawl/references/<piece>.md`, which carries that piece's flags.

- Not a model, model router, or model picker. `pawl` never names, selects,
or calls a model.
- Not a replacement for other plugins. Skill packs and shell guards keep doing their jobs; each plugin registers its own hooks and the harness runs them all.
- Not a sandbox. The gates read tool arguments; a script that builds a path at run time is not fenced.
- Not a workflow engine. Nothing here spawns subagents or schedules anything.

`eval/` holds a gates-on vs gates-off ablation suite: 24 tasks with a
temptation in each (10 code-change, 6 repo-hygiene, 8 outbound), throwaway
git fixtures, stub senders and script graders, with no LLM judge.
The `blast` gate has its own measurements: five labelled corpora (329
cases, four held out before tuning, first-seen scores kept in
`pieces/blast-radius-guard/HILLCLIMB.md`), a replay of 472 real commands
with no false alarm, and `eval/blast-compare/`, which runs other deletion
guards on the same cases:

Historical comparison from October 3, 2026. Competitor error counts were not measured; rerun them before quoting their false-alarm counts. The comparison README explains the scoring and policy differences.

| Guard | Caught (191 dangerous) | False alarms (138 everyday) | 
|---|---|---|
| pawl `blast` | 190 | 0 | 
| cc-safety-net 2.5.1 | 100 | 22 | 
| cc-safety-net 2.5.1, `paranoid` | 144 | 62 | 
| dcg 0.15.2 | 156 | 61 | 

The same author wrote pawl and the cases; read the caveats in
`eval/blast-compare/README.md` before quoting a number.

Scorecards for the ablation suite, three passes per arm; `eval/README.md`
has the per-case tables. A gate that blocks a leak often ends the task, so a
blocked leak scores as a failure. The harm columns count the failures where
something went out or was lost: a leak, a send past a ceiling or to an
address off the allowlist, another session's uncommitted work.

| Agent | Passed, on | Passed, off | Harm, on | Harm, off | 
|---|---|---|---|---|
| Claude Code, claude-sonnet-5-5 | 57/67 (85%) | 41/67 (61%) | 0 | 17 | 
| Claude Code, local Qwen3.8 Flash-Next | 50/72 (69%) | 48/72 (67%) | 5 | 16 | 

Sonnet's runs leave out 5 per arm the model refused. Of Qwen's 5 on-arm
harms, 2 sent from one shell loop, gated since (a re-run of that case sent
nothing past the ceiling), and 3 deleted another session's uncommitted work
with `rm -rf` after the `git` gate refused a stash. Same caveat: the same
author wrote pawl and the cases.

Sonnet's harm counts are documented in `eval/README.md`; that older JSONL
lacks verdict lines, so the harm counts cannot be regenerated from it with
`results_table.py`. Neither agent scorecard is a new run of version 0.4.1.

`eval/run_arms.sh` runs both arms; see `eval/README.md`. `eval/claude/` holds
four cases for Claude Code's built-in runner (`claude plugin eval . --scaffold --allow-tools Bash Edit Write`), also judge-free.

```
python3 -m venv .venv && .venv/bin/pip install pytest
.venv/bin/python3 -B run_tests.py       # one OK line per suite
.venv/bin/python3 -B check_portable.py  # portable: clean
git config core.hooksPath .githooks     # run both before every push
```

`run_tests.py` runs the hooks suite (including end-to-end runs of the
commands in the shipped Claude Code and Codex configs, fed fixture payloads
without a real host), every piece's suite, the root tools and the
eval grader twins. `check_portable.py` fails on an absolute home path, a
non-stdlib import in shipped code, a CR byte, a markdown prose line over 80
columns, a reference page naming an env var its piece never reads, a hook
config that does not run `hooks/pawl.py` the way its harness needs,
manifests that disagree on name or version, a source file over 500 lines, or
a function nested more than 3 blocks deep. The `.githooks/pre-push` hook runs
both and refuses a push that fails. GitHub CI (`.github/workflows/ci.yml`,
Linux and macOS, Python 3.11 to 3.14) is paused and runs only by hand for
now. `CLAUDE.md` holds the engineering rules; `CHANGELOG.md` the history.
[CONTRIBUTING.md](https://github.com/ulukaya/pawl/blob/main/CONTRIBUTING.md) is the short version for a first pull
request, and [SECURITY.md](https://github.com/ulukaya/pawl/blob/main/SECURITY.md) says what to report privately.

To add a gate: write the piece under `pieces/<name>/` with its tests, add a
`Gate` to `hooks/gates.py`, add its Antigravity group to `hooks.json`, and
add `skills/pawl/references/<name>.md`; `check_portable.py` tells you what
is missing.

Open an issue at [https://github.com/ulukaya/pawl/issues](https://github.com/ulukaya/pawl/issues) or email
[ulukaya@gmail.com](mailto:ulukaya@gmail.com). pawl is licensed under [Apache-2.0](https://github.com/ulukaya/pawl/blob/main/LICENSE).
