We audited 8,601 law firm websites with HTTP Archive. Three numbers were wrong until we checked by hand. Advocentral analyzed 8,601 US law firm homepages from the September 2026 HTTP Archive mobile crawl and found that three crawl-derived figures were wrong until verified by hand. The robots_txt metric's 6.6% AI-crawler Disallow rate overstated blocking — parsing the 716 named files showed only about 1% of sites block an AI crawler site-wide, with most rules coming from stock Squarespace templates — while the 28.5% unlabelled-form-input figure was dropped after Lighthouse's label audit failed on just 6%, and 23% of sites flagged as tracking-free actually loaded Google Ads or Meta Pixel tags. The team also fetched /llms.txt directly, finding roughly a third of firm sites have one, mostly auto-generated by Wix (99%) and Duda (96%). We wanted to know what the websites of US law firms actually load, so we pulled 8,601 firm homepages out of the public HTTP Archive https://httparchive.org/ crawl and analyzed them. The full results are in our 2026 Law Firm Website Report https://www.advocentral.com/blog/law-firm-website-report-2026 . This post is about the method, and about three places where the crawl data gave us a wrong answer. httparchive.crawl.pages September 2026, mobile, root pages . 9,269 matched; 8,601 were left after removing courts, bar associations, directories and non-US sites. technologies Wappalyzer detections , custom metrics.robots txt , custom metrics.structured data , custom metrics.a11y and summary . lighthouse column across the whole table: the dry run for that one column came back at 17 TiB. The robots txt custom metric counts rules per user agent. 6.6% of sites had a Disallow rule for GPTBot, ClaudeBot or another AI crawler, so that looked like the headline. It only records that a rule exists, not which path it covers. We fetched and parsed the 716 files that named an AI crawler. About 1% of all sites block one from the whole site. Most of the rest were the stock Squarespace robots.txt, which lists AI crawlers in the same group as and only hides system paths like /config and /search . Lesson: a rule count is not a block. Parse the file. The a11y custom metric includes an accessibility tree for form controls. Counting inputs with an empty accessible name gave 28.5% of sites. Then we ran Lighthouse on 1,954 of the same homepages through the PageSpeed Insights API. Its label audit failed on 6%. The likely cause is hidden fields, such as the reCAPTCHA response textarea, being counted as unlabelled inputs. We dropped the 28.5% figure. What Lighthouse did show: 68% of homepages fail color contrast, 54% have links with no accessible name, and the sites using an accessibility overlay widget scored no better than the rest median 85 against 89 . This one was too low, not too high. We loaded 227 homepages in a real browser and compared them with the crawl's Wappalyzer flags. 23% of the sites flagged as clean did load a tracking script, mostly Google Ads tags and Meta Pixels injected through a tag manager. Detection of consent tools was off in the other direction: 17% of the sites flagged as having none did show a cookie banner. Lesson: technology detection gives a floor. If a number matters, check a sample by hand and publish the error rate with it. HTTP Archive doesn't request /llms.txt . We fetched it ourselves: about a third of law firm sites have one. Wix and Duda generate it for nearly every site 99% and 96% , and on WordPress it usually comes from an SEO plugin. Most firms probably don't know the file is there. The report's charts are plain HTML and CSS with a 2 KB script: no chart library, bars that render without JavaScript, and a tile map for the state data. It seemed wrong to ship a heavy page about heavy pages. Full numbers, method and limits: The 2026 Law Firm Website Report https://www.advocentral.com/blog/law-firm-website-report-2026 . Disclosure: Advocentral builds websites for law firms. The report names no individual firm.