Coding Agent Proof Spirals A developer who works with coding agents on long autonomous tasks has documented an anti-pattern called a "proof spiral," in which an agent writes checks of its checks — building receipt checkers, fingerprints and provenance chains — until delivery stalls while logs continue to look rigorous. The write-up attributes the behavior to workflow rules such as "always verify independently," "fix every finding" and "persist until done," which it says together prevent completion, and it prescribes a single stopping rule plus the question "What decision does this proof change?" The author published the guidance as a drop-in skill, proof-spiral.md, and also introduced ThinkThen, a tool with ten functions including decide, choose, tag and score that runs across 25 surfaces and returns true (exit 0), false (exit 1) or not sure (exit 3). Coding Agent Proof Spirals coding agentsagent designskills I work with coding agents on long tasks where they plan and develop on their own. Most of the time it works. Every so often an agent falls into an anti-pattern I call a proof spiral. It starts writing checks of its checks. What it looks like The agent builds a feature and writes a check. The check needs a record. The record needs a review. The review finds a gap in the check, so the agent builds a better checker. Each step looks reasonable. Together they stall delivery, and the logs look rigorous the whole time. These are the signs I watch for: - The same check reruns while nothing it depends on has changed. - New gates appear after the existing gates pass. - Review findings never connect to a product failure. - The agent builds tools to check its checks: receipt checkers, fingerprints, provenance chains. - I ask for status, and the agent describes proofs. It doesn’t tell me what works. Watch the status lines too. When they fill with words like frozen, sealed, byte for byte and exact receipt, and keep ending in “remains pending”, I go look. Why it happens It’s tempting to blame the model. Look at the workflow first. Workflow rules often make defects and violations explicit. They rarely say how long delivery should take or how much uncertainty is fine to leave behind. Take three reasonable rules: “always verify independently”, “fix every finding” and “persist until done”. Together they prevent completion. A model that follows instructions well obeys all three. Strong instruction following becomes the liability. Rules also pile up. Each bad run earns a new rule, and the old rule stays. Checking can also stand in for access. An agent that can’t reach the real environment builds mocks and checkers instead. More local checks can’t settle an external question. How to stop it Start with one question: What decision does this proof change? If the answer is none, the proof is ceremony. In the spirals I’ve seen, this question would have stopped them earliest. Then fix the workflow: - Use one stopping rule. Complete the agreed checks and investigate concrete failures. Repeat or widen a check only when something relevant changed, or when the check can resolve a named uncertainty. Otherwise deliver. - Replace conflicting rules. Don’t layer an override on top. A temporary exception leaves the old rules in place. - Size checks by consequence. A one-line authorization change can be critical. A wording fix may need only a look. Diff size tells you little. - Name what stays unknown. When the available checks can’t resolve something, record the limit and deliver the finished work. If an agreed acceptance criterion stays unmet, call it a release blocker. Don’t claim readiness, and don’t invent substitute proof. Heavy checking isn’t always a spiral. Releases, migrations and security changes can need it. Keep the checks that earn their place: behavior tests, one review, the release rehearsal, and checks on money, data, credentials and users. The skill I wrote this up as a skill you can drop into your own agent setup: proof-spiral.md https://gist.github.com/imaurer/79c1b9f3bef3fea0f4c99f9342905a5c . It covers how to notice, diagnose, address and prevent a spiral. It also warns about itself. The skill must not become one more layer of rules. Also: introducing ThinkThen I recently introduced ThinkThen. It answers typed questions about text and returns true , false , a label or a number. A failed call never looks like an answer. Use it to gate a script, label records, or grade answers in an eval. thinkthen decide 'Does the customer ask for money back?' < message.txt The command prints true and exits 0 for yes, 1 for no and 3 for not sure. A shell if can branch on it. It has ten functions, such as decide , choose , tag and score . They run across 25 surfaces: the command line, a Rust crate, language bindings from Python to COBOL, and SQL extensions for DuckDB, SQLite and PostgreSQL. Read the docs at thinkthen.dev https://thinkthen.dev . The code lives at github.com/botassembly/thinkthen https://github.com/botassembly/thinkthen .