A test system that can say "I don't know" is worth more than one that says "passed" Engineer Derek Wang argues that AI coding harnesses should be built as Popperian refutation machines rather than pass/fail gates, citing a full regression suite that runs in ten seconds. His `fulltest` tool classifies failures against a growing ledger of known failure portraits, treating unmatched failures as "unknown" and therefore the most valuable signal, and grows that ledger from incident logs, diagnosed issues, and git fix commits. AI Harness Engineering · Essay Eleven · Derek Wang Karl Popper spent the middle of the twentieth century dismantling a three-hundred-year-old story about science. Since Galileo, the comfortable version was that knowledge piles up: prove one thing today, one thing tomorrow, truth accumulating. Popper said the opposite. Science does not advance by being right more often; it advances by wrong answers being torn out of the game in broad daylight. A beautiful theory is worth exactly what it survives — not what it explains. Every time someone reaches for it and knocks it flat, that failure still did the work, because it removed one wrong possibility from the board. He called that guess-and-refute, and it is the sharpest one-sentence mission statement ever written for testing. To harness an AI — steer what it is allowed to generate, restrain where that force may land — is to build that same disagreeable machine for code. Not a machine that proves the code is right. A machine that keeps trying to prove it wrong, and keeps a ledger of every failure it has ever seen. That ledger is the whole argument of this essay, and it sits where most growing AI systems secretly have nothing at all. Every team's testing starts the same ugly way. An agent touches a contract, squints at the output, declares it good. The rule that kept us honest before we had a name for it: no baseline, no accountability. If you have never counted your failures, you do not even know how much you owe. Our full regression eventually ran in ten seconds ORIGINAL DATA . Ten seconds reads as free — and that is exactly why it matters. Ten seconds is fast enough to skip, and skipping verification is how confidence quietly turns into a lie. Speed and safety are the same lever here. The faster the feedback, the bolder the correct change and the more timid the wrong one. Why frame it as a ledger instead of a gate? Because a gate only stops bad work; a bridge lets good work travel. Trustworthy feedback changes what the agent dares to do. It can change a contract, refactor a layer, and know in the same breath whether it broke the world — instead of firing into the dark and betting it does not blow up. The hill most teams die on: they decide self-growth means "fix a bug, write a test for it." That is the skin, not the bone. One more test is one more isolated check. What actually makes a system evolve is the moment it starts to recognize failure — whether this one has come before, whether it is an old wound or a fresh hole, whether it is fair to charge it to this change at all. A test system that only says "pass / fail" is, in Popper's terms, degenerating science: it counts its green lights and never asks whether the error list got longer. So fulltest grew three things a plain runner never thinks about. Those three are the only part worth calling self-growth. One: classify before you judge. fulltest keeps a ledger — a portrait for each class of failure: the shape it takes, a one-line root cause, a repair recipe that has worked before. Every full run ends matched against that ledger. Matching a portrait is a known failure — a debt we already account for, old rule applied, move on. Failing to match is the gold: an unknown failure, filed separately, earmarked as raw material. The deliberately counterintuitive part: the scariest sentence in the ledger is not "a known failure came back," it is "there is one we have never seen." Known is accounted debt; unknown is a hole you did not know existed. Route one move, and testing stops being "check whether the machine is good" and becomes "measure how complete our map of failure is." Two: the ledger is grown, not written. Nobody writes it from a whiteboard. It grows from real incidents, fed from three mouths: the unknown-failure log every unmatched failure is a candidate , the diagnosed issue once a bug is traced to root and fixed, the whole cause→remedy experience goes in , and the fix commits in git any commit that says what it repaired . Growth is housekeeping between runs, owned by the organization, not any single execution. A new project does not start by re-chewing every mistake; it arrives already carrying the scars of the projects before it. Three: the baseline climbs on its own. Every run compares this count against the recorded baseline. Stronger, and the new number becomes the baseline and the version bumps — the floor just rose, you cannot silently backslide below it. Weaker, and a regression alert names the class coming back: you did not just fail to improve, you slipped. Unchanged and clean is the quietest good news: everything passed, the floor holds its post. The deepest insight of the essay hides here: a baseline is not a fixed pass-line. It is how you harden every improvement into a floor you cannot quietly walk back under. The system is not standing guard; it is driving wedges into the slope as you climb. To go backward, you first have to pry the wedges out. Cold water, always. The real disease of every testing regime is falling in love with the number. Fifty tests, six hundred, six thousand — none of it means anything if the tests measure the wrong thing, sit on duplicated noise, or never touch the paths where failures actually live. A wall of a thousand sentries who never see combat is just a very expensive garden fence. The measure that matters is not how many tests you have. It is how many past real incidents the suite would have caught . If adding a scenario never changes the answer to that question, you are not growing a wall. You are growing a museum. Dropped into concrete: one ten-second run grew into fifty scenarios organized into six suites ORIGINAL DATA — contracts, idempotency, retrieval, regression, resilience, end-to-end. Six directions of attack, each lane an archive of an earlier mistake. It paid for itself on a release where 22 hidden HTTP-500 errors sat in the deep path ORIGINAL DATA ; the full suite caught every one before shipping. Twenty-two ways the system was quietly broken, all surfaced before a user saw them. That is not the AI being clever. That is the ledger doing what it exists to do — and it is why those twenty-two will never reach production a second time. Before this becomes a hymn, one admission a serious testing essay owes you: a test suite can only hold the failures you were smart enough to imagine and articulate. Every row in the ledger exists because a human — or a model — first thought of a failure and then wrote a guard against it. The most expensive failures in engineering are exactly the ones nobody thought of. So one day a release broke with every lane green. The six suites passed, the integration test was 62/62 ORIGINAL DATA , no red light anywhere. It still broke in production — on a boundary nobody had turned into a test. A third-party endpoint returned a malformed shape in a narrow time window, and the error multiplied up the whole chain into an incident. When the team rushed to add a test afterward, they could not even reproduce it reliably; the anomaly only appeared under a specific data distribution the test environment could not fabricate. The failure was not "we did not write enough tests." You can never out-write the failures you have not imagined. That is where the layer beside fulltest earns its keep: failure-mode analysis . It does not ask "will this failure come?" It asks "if it comes, does the system survive?" A boundary anomaly we cannot test for becomes a question of whether the parser can tolerate and degrade, keep the core path alive, and log the anomaly as a new ledger entry before it becomes a crisis. So the two tools split the world cleanly. fulltest turns the failures it can see into a ledger that grows more complete. Failure-mode analysis leaves a fire escape in every corner the tests cannot see. One accountants the known; the other defuses the unknown. Both write into the same debt book — and a system that learns to recognize its failures, and to survive the ones it cannot, is the only kind that stops being fragile.