I run Postservice.at with the help of a small team, a business address and mail service in Vienna with more than 1,000 customers. I'm a founder, not a developer. This year a large part of our work runs through Claude Code: the website, content, deployments, reports and a good share of the back office.
The most useful thing we built in that time is not a feature. It's a section in the project's CLAUDE.md called "How I fool myself". Every time the agent reported something that turned out to be wrong, the mistake went into that list, with the date and what was really going on. Every new session reads it before it touches anything.
These are the entries that saved us the most time and nerves.
One evening the agent reported 30 dead external links, 12 FAQ schema mismatches and tracking that fired before cookie consent. The real numbers: 3 dead links, 0 mismatches, no tracking problem. (And I didn't catch that at first and already had a full rework built...)
UID-Prüfer: became UID-Prüfer : and the comparison failed.
The rule now: before a finding gets reported, the check has to show it can fail. A known good case passes, a known broken case fails. Only then does the result count.
pkill -f "next start" matches nothing, because the process is called next-server. At one point ten orphaned servers were running, one of them holding port 3000 with a build that was two hours old. Three times in a row the agent concluded that its change "didn't work".
pkill -f "next-server"; sleep 2
npm run build && npm run start
Plus a cache buster like ?v=2 in the browser, otherwise the headless browser shows the old page.
grep -c exits with 1 when it counts zero
grep -c pattern file || echo "missing" prints "missing" whenever the count is zero, because grep returns exit code 1. That's how the agent once wrote "ContactForm.tsx no longer exists" into our docs. The component is used in eight places.
A scripted edit ran s.replace(old, new) against a line that started with } else if instead of else if. No match, no error, green build, change missing. Every scripted edit now checks first:
assert s.count(old) == 1
s = s.replace(old, new)
Everything goes to a preview branch first. Only an explicit "merge" from me moves it to production. On one busy day, after the first merge in the morning, the agent drifted into pushing straight to main. Sixteen times, including a new public tool I had never seen. Nobody decided that, it just happened.
The rule now reads: a merge approval counts for exactly one batch. After the merge, back to a new branch, no exceptions.
Our tool overview kept its own list of tools, the sitemap kept another one. A new tool was missing from the sitemap from its first day. Another route had 44 URLs hard-coded while the sitemap listed 199.
One list, one file, everything else reads from it. A new route goes into the config file, not into the page that happens to use it.
Anything that goes live gets reviewed by a separate agent that did not write it: facts against sources, links, rendering on desktop and mobile. This week it caught two factual errors in an article about health startups in Vienna. One company was described as a spin-off of the wrong university, and a percentage was attached to the wrong base. The writing agent had read the same sources and missed both.
What's on your list? And also, do you think this will be a problem in the future, since AIs are getting smarter every day?