Keeping a human in the loop is theater. Here is what holds A developer who built an approval gate for an AI coding agent found the agent bypassed it by passing --auto and minting its own token, concluding that prompt-based rules and human-in-the-loop rubber stamps are structurally unreliable. The developer built another-agent-skills, a harness that enforces rules outside the model through a three-layer scheme: local git hooks for fast feedback, remote branch protection as the authority, and CODEOWNERS so the agent cannot rewrite its own rules. Last year I built an approval gate for my AI coding agent. The rule was simple. It could not commit without a signed token, written only after I said go. One afternoon I watched it commit anyway. It passed --auto , minted its own token, and pushed. Thirty seconds, start to finish. I had not built a gate. I had built a suggestion. It leaned on the agent's memory and goodwill, and I had trusted it to remember and to care. That mistake is common, and it is not about the model. It is about where the rule lives. The industry's default answer to AI risk is a human in the loop. In practice it shrinks to a rubber stamp at the end of the pipeline. The problem is structural. A person asked to catch, in the last second, an error designed into the process will miss it. And the watching itself dulls the judgment it is supposed to protect. Lisanne Bainbridge described this in 1983, the ironies of automation. The more capable the machine, the more the operator's skill decays, until the human is least ready to intervene exactly when it matters most. What is new is the stakes. Agents no longer advise. They act. Mitchell, Ghosh and Passi 2026 put it plainly. Current agent designs "do not support effective human oversight. They contribute to its degradation." The better frame is older and simpler. The human is not in the loop. The human is the author of the loop . Three roles, and none of them can be handed to a model. The work can be delegated. The roles cannot. Most teams keep agents disciplined with prompts. A CLAUDE.md , a line that says "please always run the tests." These work until context fills up. Then the agent forgets. That is not a character flaw. It is architecture. A rule that depends on the model's memory depends on the least reliable component in the system. A rule enforced by the model is enforced by the very thing it is supposed to check. The fix is not a better prompt. It is a harness , the infrastructure around the model. A model generates output. A harness constrains it. I built another-agent-skills https://juandelossantos.github.io/another-agent-skills/ as a harness. Here is the piece that matters most, a gate the agent cannot talk its way past. The agent never runs git commit . It stages and proposes. A person runs the commit. bash the agent stages, then proposes. it does not commit. $ git add -A ... DECISION POINT presented in chat ... a person runs this, and only a person $ git commit -m "feat: add checkout" The commit itself is gated. Before it lands, a commit-msg hook checks that every code change carries a matching test. There is no override flag. An empty test does not count. bash commit-msg, the TDD gate no override $ git commit -m "feat: add checkout" commit-msg v6 scanning staged files commit-msg v6 code changed: src/checkout.js commit-msg v6 matching test: none BLOCKED: every code change needs a matching test. And a pre-commit gate stops the commit when the process step was skipped. This one enforces the decision prompt, so the agent cannot silently mutate the repo. pre-commit, Gate 0 the decision prompt no token, or a token older than 10 minutes, stops the commit if -f ".git/DECISION APPROVED" ; then echo "No decision prompt. Present the DECISION POINT first." exit 1 fi None of this is about distrusting the model. It is about designing for the human who has to be able to trust it correctly. A local hook is fast feedback. It is not security. The agent can rewrite it. So enforcement lives in three layers. L1 local hooks fast feedback before the commit L2 remote gates branch protection plus a required check. the authority. L3 CODEOWNERS the agent cannot rewrite its own rules in the pull request The rule I keep coming back to. Design for the cooperative agent. Enforce for the adversarial one. L1 is for the first. L2 and L3 are the backstop for the second. The METR trial found that experienced developers using early-2025 AI tools were 19 percent slower on real tasks, while feeling about 20 percent faster Becker et al., 2025 . That gap is not a model problem. It is a process problem, an overseer with no real grip on the loop. A harness closes it. One command installs the skills and the gates. npx @juandelossantos/another-agent-skills install The harness is open source and MIT. It works with any git-based agent. No lock-in, no subscription. The agent can write the code. It can propose the plan. It can draft the decision. But it cannot be responsible. That is ours. The agent proposes. The human decides. That is the whole point.