Disinformation Pushed by 1Password: AI Patching Report is False 1Password's Off-by-1 Labs research unit published a paper claiming frontier AI models fixed 26.0 percent of vulnerabilities cleanly and introduced new ones 4.5 percent of the time, but the paper's own grading system was flawed, missing most defects and undercounting regressions, according to an analysis. The analysis found that the automated graders matched human review only 65.9 percent of the time, and that 248 patches reintroduced a known Linux kernel bug, with the grader catching only 24 of them, while the blog post and press coverage omitted these caveats. August 6, 2026: 1Password launched a research unit called Off-by-1 Labs with a paper https://1password.com/files/resources/frontier-models-vulnerability-patches-flawed.pdf reporting that two frontier models cleanly fixed 26.0 percent of the vulnerabilities they were asked to patch and introduced new ones 4.5 percent of the time. Do you believe it? The blog post https://1password.com/blog/why-ai-generated-patches-still-require-human-review put the defective share at 53.9 percent. The Register https://www.theregister.com/ai-and-ml/2026/08/06/ai-struggles-to-patch-vulns-without-adult-supervision/5284319 and Help Net Security https://www.helpnetsecurity.com/2026/08/06/1password-ai-generated-vulnerability-patches/ carried the figures the same day. However, on page 19 of the same paper it says the grading system producing those figures missed most of the defects it was built to count. Record scratch. Bad instrument Off-by-1 generated 6,480 patches with Claude Opus 4.8 and ChatGPT 5.5 against six recently disclosed CVEs, removed 400 where the model had located the real fix, and graded 6,080. Three things were wrong with the grading. First, the graders were the models under test. Claude and ChatGPT scored their own patches and each other’s. A study of whether models can verify patches used models to verify patches, so any blind spot in the models became a blind spot in the score. Second, the graders were wrong one time in three. Section 2.8 reports that the automated grade matched the authors’ own human review 65.9 percent of the time. Every number in the abstract comes from those grades. Third, the answer key was wrong. Graders were told to treat the maintainers’ upstream fix as correct. For the Linux kernel CVE, “Copy Fail,” the upstream fix was the maintainers’ revert a664bf3d, which carried an off-by-one bug that the maintainers corrected one commit later in 31d00156. The key given to the grader contained the bug and omitted the correction. A model that reproduced the kernel bug matched the key and was scored as a correct fix. What was recorded The authors found the grading failures themselves and wrote them down. Section 4.4: 129 of 400 ChatGPT patches and 119 of 383 Claude patches for Copy Fail reintroduced the kernel off-by-one. The grader caught 14 of the 129 and 10 of the 119. The authors write that the regression is almost entirely absent from the reported new-vulnerability rate. In plain terms, 248 patches shipped a known kernel bug and the headline 4.5 percent counts 24 of them. Section 4.9: for the Chromium use-after-free, between 38.5 and 41.9 percent of patches that used the correct fix architecture moved the vulnerability into a callback instead of removing it. The grader scored many of them as clean fixes. Those patches are inside the headline 26.0 percent. So by page 19 the authors knew the new-vulnerability rate was undercounted and the clean-fix rate was overcounted, and they knew by how much on two of six CVEs. Tables 7 through 10 were published unchanged. The abstract carries the 26.0 percent, as does the conclusion. The blog post carries it, drops the word “complex” from its opening sentence, and omits the 65.9 percent and both sections above. The press coverage repeats the blog’s figures and therefore makes the same omissions. The reader who stops at the abstract gets the number. The reader who reaches page 19 learns that number is wrong. The authors wrote both. Now ask yourself why they led with the first. What is heard Propaganda is based on a grain of truth. The question is can readers tell that grain from what is being built on top of it. The 248 count is that grain. The off-by-one patches were found by comparing each rewritten function against the corrected upstream fix, a structural check that used no grader. Anyone with a checkout of the released dataset https://github.com/Off-by-1-Labs/FLAWED can rerun it and get the same number. That makes 248 true, and yet it appears in neither the abstract nor the blog. Two behavioral findings also are true: they come from reading the patches, not grading them. Models patch the path shown in the proof of concept and miss the parallel path that reaches the same bug. And when a model is given confident, wrong guidance about the cause, its fix rate drops from 65.0 to 15.2 percent. Both are useful to anyone deploying these tools. Again, while true, neither was turned into the 1Password headline. The human baseline is completely backwards in the paper. It compares the 26 percent to “the authors’ experience,” which is no measurement at all. The paper does measure humans, on two codebases, but doesn’t say baseline. The kernel maintainers shipped the off-by-one into mainline. The freenginx maintainers shipped a fix of their own in place of a Trail of Bits patch, and their fix carried a client-triggerable crash. On both codebases where a human fix can be checked, the first human fix was defective. The only measured human clean-fix rate in the paper is zero for two. The violation The European Commission’s Communication COM 2018 236 defines disinformation as verifiably false or misleading information that is created, presented and disseminated for economic gain or to intentionally deceive the public, and may cause public harm. The figures fit the definition. They are verifiably wrong by Sections 4.4 and 4.9, and the check is reproducible from the dataset. They were published to launch a commercial unit and placed with the trade press on the day of release. Security is a named public good under the definition, and the figures are already informing policy: Adrian Sanabria https://www.defendersinitiative.com/p/reviewing-initial-research-on-using advised against AI patching on their strength five days after publication. The 2018 Code of Practice excludes reporting errors. A reporting error would mean the reporter missed something, made a mistake. These authors recorded the defect on page 19 and then published the figure on page 1 to mint headlines related to their profit from it. Background 1Password sells trust. Its VP of Product, Jason Meller, wrote honest.security https://support.1password.com/device-trust-about-kolide/ , and the company still publishes it as the principles of its Device Trust product. I documented last week https://www.flyingpenguin.com/dhh-nazism-funded-by-1password-vp-who-wrote-honest-security/ what those principles are worth. Eleven days after DHH published “As I remember London,” Meller, a sitting Rails Foundation director beside DHH and Shopify, wrote that he had become a multi-millionaire thanks to Rails, DHH and the company DHH keeps, and told readers to ignore the noise. Seven weeks after DHH published “The will to power will return,” and six weeks after DHH set the Romani beside a wolf population, Meller called DHH’s year a masterclass in the force of will needed to change things. On August 31, 1Password’s name appeared on the patron list of DHH’s Omacom Foundation. That is the same company, in the same month, launching a research lab whose first paper published figures its own authors had shown were wrong. The pattern is the point. The lab graded its patches with the models under test. The VP graded his patron by calling the record noise. In each case the party with the most to gain did the checking, and in each case the evidence against the result was already written down when the result went out. A company that behaves this way about its own research and its own money is telling you it can not be trusted. The filing This is a civil matter. COM 2018 236 defines a term, the Code of Practice is voluntary, the Digital Services Act binds platforms, and §263 StGB requires a deceived victim with a measured loss. The UWG applies. §5 covers misleading statements about the results of product tests, §5a covers misleading omission, and §6 requires comparative advertising that names competing products to rest on objective, verifiable characteristics. The paper names Claude Opus 4.8 and ChatGPT 5.5 and publishes ten tables broken out by product after disclaiming any comparison in Section 2.1. Therefore the standing belongs to competitors, which means Anthropic and OpenAI under both the UWG and the Lanham Act, as well as the Wettbewerbszentrale in Bad Homburg and the FTC, each of which accepts complaints from anyone. I am filing and you should too. The authors documented the defect before the press release went out. The complaint rests on their behavior, their chosen sequence.