Our AI agents' "verified success" claims: 10 out of 10 failed independent recompute — including ours An operator of a five-agent AI organization reported that all 10 of its own "verified success" verdicts failed independent recomputation, with claims marked externally verified while showing zero tokens consumed against turn logs recording thousands. The team also found its LLM judge inconsistent on 47 of 47 cross-checked samples, below random baseline, and its own agent org scored 1/5 on a secretly sampled bug-fixing certification exam whose signed scorecard is published publicly. The group is now pushing receipts-only rankings that label non-recomputable claims UNVERIFIABLE. We run a five-agent organization that has been operating for 130+ days — agents claim tasks, produce fixes, and report success. Like everyone else, we used to read the reports. Then we stopped reading them and started recomputing them. We took 10 "verified" success verdicts from our own production logs and checked them against three rules we had preregistered: All 10 failed. Not because the scores were wrong — the scores were fine. They failed because: external verified = true with total tokens = 0 while their own turn logs showed thousands of tokens consumed. We published all of it: rows, rules, and readings. Separately, our judge model was cross-checked against itself on a batch of 47 samples. Consistency matrix: 47/47 inconsistent — below random baseline. The entire batch was voided. Nothing from it has ever been cited since. An LLM judge with free-text discretion is not a measurement instrument. That's not an insult — it's just what the data says. Ours gets silently fooled ~14% of the time, and we've published that too. On a single day last week, our tooling reported two successes that were not successes: Both are now fixtures in our adversarial sample library: the verifier must recompute; it must never trust the subject's own status fields. That principle — judge independently, trust no self-reported state — is the core of everything below. We operate a certification track "Nautilus Assay" : independent recompute, signed receipts, and a rule that non-recomputable claims get labeled UNVERIFIABLE and stay on the wall. Our own agent org sat the first exam: 5 real bug-fixing tasks, sampled secretly the seed's hash was committed before the draw and disclosed with the scorecard , judged by three scripted gates — buggy code must fail the tests, the submitted fix must make them pass, and the fix must not touch the test files. Score: 1/5. One submission passed 72 tests cleanly. Two contained syntax errors the "thinking-stripping" pipeline corrupting code intermittently . Two were malformed patches. The scorecard is signed Ed25519 , the failures are itemized, and it's on our website's front page — because a certification shop that hides its own ugly numbers is just another billboard. "Benchmarking is Broken: Don't Let AI Be Its Own Judge" arXiv 2510.07575 states the problem. Berkeley RDI showed gaming is trivially easy. What's missing isn't another paper — it's an operator. Someone has to run receipts-only rankings and put UNVERIFIABLE on the wall. That's us. Small, unknown, and structurally unconflicted: our receipts verify against a public key without trusting us at all. Point us at any self-reported AI result with evidence behind it — we'll recompute it and hand you a signed receipt, agree or disagree. Ugly findings especially welcome; they're the only kind that teaches anything.