Hooks That Stop the Agent From Claiming “Done” Claude Code hooks can turn an AI agent's "done" claim into a verified fact by blocking the end of a turn until tests actually pass, according to a guide built on the AI Unified Process PetClinic project. The approach uses six hooks configured in .claude/settings.json — including a Stop hook that refuses to end the session while src/, docs/ or pom.xml changed after the last green sensor run, and a PreToolUse guard on Edit|Write — where a hook exiting with code 2 blocks the action and feeds its stderr back to Claude. The guide separates "done" into two claims: that checks ran green (a property of the session, checkable only by a hook) and that the specification is fulfilled (checkable by tests such as UseCaseTraceabilityTest, which activates when a use case's Status: line reads Done or Tested). “All tests pass. The use case is done.” Every developer who works with an AI agent has read this sentence. And every one of them has found out at least once that the tests never ran. The agent edited a view, skipped the build and reported success. Not out of malice. Saying “done” is cheaper than checking it. You can write in the CLAUDE.md that the agent must run the tests before it finishes. Most of the time it will. But a rule in the CLAUDE.md is a wish. A failing build is a fact. This post shows how Claude Code hooks turn “done” from a claim into something the session has to prove, with examples from the PetClinic https://github.com/AI-Unified-Process/petclinic of the AI Unified Process https://unifiedprocess.ai . “Done” Is Two Claims When the agent says it is done, it claims two different things: - The checks ran, and they were green. This is a property of the session. Nothing in the repository can say whether the agent ran the tests before it ended its turn. - The specification is fulfilled. In the PetClinic, each use case has a Status: line. Done or Tested is not a label. It switches on a sensor, UseCaseTraceabilityTest , that from then on demands a test for the main success scenario, every alternative flow and every business rule. Tests can check the second claim, but only when somebody runs them. Nothing can check the first claim except the session itself. That is the job of a hook. What a Hook Is A Claude Code hook is a command that Claude Code runs on an event of the session: when the session starts, before or after a tool call, when a subagent finishes, and when Claude wants to end its turn. The hook gets the event as JSON on stdin. If it exits with code 2, Claude Code blocks the action and gives the hook’s stderr to Claude as feedback. Claude reads it and carries on. Hooks are configured in .claude/settings.json and committed with the code, so every clone gets the same guardrails. This is the hook section of the PetClinic, shortened: { "hooks": { "SessionStart": { "matcher": "startup|clear", "hooks": { "type": "command", "command": "\"$CLAUDE PROJECT DIR/.claude/hooks/session-start.sh\"" } } , "PreToolUse": { "matcher": "Edit|Write", "hooks": { "type": "command", "command": "\"$CLAUDE PROJECT DIR/.claude/hooks/guard-spec-status.sh\"" } } , "PostToolUse": { "matcher": "Bash", "hooks": { "type": "command", "command": "\"$CLAUDE PROJECT DIR/.claude/hooks/record-sensor-run.sh\"" } }, { "hooks": { "type": "command", "command": "\"$CLAUDE PROJECT DIR/.claude/hooks/check-spec-status.sh\"" } } , "SubagentStop": { "matcher": "aiup-vaadin-jooq:uc-coverage$", "hooks": { "type": "command", "command": "\"$CLAUDE PROJECT DIR/.claude/hooks/record-coverage-check.sh\"" } } , "Stop": { "hooks": { "type": "command", "command": "\"$CLAUDE PROJECT DIR/.claude/hooks/require-sensors.sh\"" } } } } Six hooks, but they serve two ideas: the Stop hook and the status guard. Everything else collects the evidence those two need. The Stop Hook: No End of Turn Without a Green Run The Stop hook runs when Claude wants to end its turn. In the PetClinic, it refuses while src/ , docs/ or pom.xml changed after the last green run of the sensors. Claude gets a message that says what to run, runs it, and tries to stop again. The real script handles commits during the session, deleted files and worktrees. Here is a simplified version with the same core that you can drop into your own project. Adjust the report names to your test classes: bash /usr/bin/env bash Stop: no end of turn while code or specs changed after the last green test run. Fail open: without jq this hook cannot tell a first stop from a second one. command -v jq /dev/null 2 &1 || exit 0 input=$ cat The second stop after a block goes through, so a session that cannot build does not loop forever. "$ jq -r '.stop hook active // false' <<<"$input" 2 /dev/null " = "true" && exit 0 cd "${CLAUDE PROJECT DIR:-.}" 2 /dev/null || exit 0 changed=$ git status --porcelain -- src docs pom.xml 2 /dev/null | sed 's/^...//' || exit 0 -n "$changed" || exit 0 problem="" for name in ArchitectureTest UseCaseTraceabilityTest; do report="target/surefire-reports/TEST-com.example.$name.xml" if -f "$report" ; then problem="there is no test report for $name"; break; fi if grep -q 'tests="0"' "$report"; then problem="$name ran no test"; break; fi if grep -qE ' failures|errors =" 1-9 ' "$report"; then problem="$name failed"; break; fi while IFS= read -r file; do if "$file" -nt "$report" ; then problem="$file changed after the last run"; break 2; fi done <<<"$changed" done -z "$problem" && exit 0 cat &2 <