We are a small family business in Spain, second generation. We service and repair car-wash equipment (gantries, tunnels, jet-wash bays) for two of the big manufacturers and for our own customers. A handful of field technicians, one office, and me running it. Nobody here is a programmer.
#
What we built
Our own ERP, in production , replacing the commercial package we used for years. It handles service calls, work reports from the technicians' phones, quotes, purchase orders, supplier invoices, bank reconciliation, payroll inputs and, since this month, the electronic invoicing regime that is becoming mandatory in Spain. Invoices are signed and chained; once issued, the database itself refuses to let anyone change them. #
A set of agents built on Claude Code and an MCP server. They read the mailbox, create service calls from manufacturer emails, draft replies, reconcile what the manufacturer approved against our work reports, prepare supplier orders, and prepare (never send) emails. #
A technical library and an "expert" the technicians can ask. The manufacturers' manuals and parts diagrams, plus our own history of faults and fixes, indexed and linked to each machine model. A technician can ask from the phone "this model does X" and get the likely causes, what we did last time on that same machine, and the page of the manual. Every answer can be marked right or wrong, and the corrections go back into the library. #
Human approval by Telegram. Anything with consequences (creating a document, changing a price, sending something) shows up on my phone with two buttons. Every decision is logged: who, when, result. #
A second model, from a different company , that reviews what the first one did and flags disagreements. #
A base of written procedures , about 1,500 of them, that the agents must consult before they are allowed to write anything. The tooling enforces it: no procedure read, no write. #
Local models on our own hardware for the cheap, repetitive parts: reading and classifying documents, embeddings for the library, a first pass at reviewing code. Judgement stays with Claude. #
Backups four times a day , encrypted, to three places, and we actually restore them to a clean machine every few weeks to prove they work.
#
What worked
Writing the procedures down. Not for the AI, for us. The AI just made it worth doing. Every time something went wrong it became a rule with a date and a reason, and the agents read those rules. #
The approval loop. I was afraid it would be a bottleneck. It is the opposite: I review twenty decisions a day from the phone in a few minutes, and nothing surprising ever happens. #
Keeping production and testing in separate copies with separate databases. The one time we built straight on production, a syntax error took the system down for a minute while technicians were on the road. Never again. #
Tests for the money logic: numbering per series, VAT, rounding to the cent, checked against the real SQL on hundreds of random documents.
#
What did not work, and what we changed
- Our first approach let the agents act and tell me afterwards. We lost a day to a batch of work reports that were "approved" with the wrong prices. Now nothing with money moves without a human click.
Long sessions. We measured it: one working session re-read hundreds of thousands of tokens on every step because it kept the whole conversation. We switched to one session per topic, with a short hand-over note, and the same work now costs a fraction. #
Letting the AI "remember" things in chat. If it is not written in a procedure or a file, it does not exist the next day. #
Local models beyond their lane. Fine for small, well-specified jobs; they derail on anything longer than a page, and giving them more context makes them worse, not better. We keep them for the mechanical work, and nothing else. #
The thing we still have not fixed: until this week our money tests did not run automatically before every change, and we had no tests for cancellations, retries or the signed chain. We added them a few days ago; they are brand new and need mileage.
#
Lessons for owners who are not programmers
You are the systems engineer. The model writes; you decide what "correct" means, and you have to be able to tell when it is wrong. 2. Make the AI read before it writes. Procedures, module docs, "what changed and why". Then make it mandatory, not polite. 3. Separate "propose" from "do". Everything that touches money, customers or production goes through a human. 4. Back up, and restore. A backup you have never restored is a hope. 5. Expect to be the one maintaining it. If a change breaks something and ten prompts do not fix it, you need to understand enough to find the problem.
#
A question for you
For those of you running invoicing or accounting logic you built this way: how do you test it? Do you have automated tests for numbering, rounding and cancellations, or do you rely on checking by hand? And how often do you actually restore a backup? Written with help from the same assistant, reviewed and edited by me.