We wanted to know what the websites of US law firms actually load, so we pulled 8,601 firm homepages out of the public HTTP Archive crawl and analyzed them. The full results are in our 2026 Law Firm Website Report. This post is about the method, and about three places where the crawl data gave us a wrong answer.
httparchive.crawl.pages (September 2026, mobile, root pages). 9,269 matched; 8,601 were left after removing courts, bar associations, directories and non-US sites.technologies (Wappalyzer detections), custom_metrics.robots_txt, custom_metrics.structured_data, custom_metrics.a11y and summary. lighthouse column across the whole table: the dry run for that one column came back at 17 TiB.
The robots_txt custom metric counts rules per user agent. 6.6% of sites had a Disallow rule for GPTBot, ClaudeBot or another AI crawler, so that looked like the headline.
It only records that a rule exists, not which path it covers. We fetched and parsed the 716 files that named an AI crawler. About 1% of all sites block one from the whole site. Most of the rest were the stock Squarespace robots.txt, which lists AI crawlers in the same group as * and only hides system paths like /config and /search.
Lesson: a rule count is not a block. Parse the file.
The a11y custom metric includes an accessibility tree for form controls. Counting inputs with an empty accessible name gave 28.5% of sites.
Then we ran Lighthouse on 1,954 of the same homepages through the PageSpeed Insights API. Its label audit failed on 6%. The likely cause is hidden fields, such as the reCAPTCHA response textarea, being counted as unlabelled inputs. We dropped the 28.5% figure.
What Lighthouse did show: 68% of homepages fail color contrast, 54% have links with no accessible name, and the sites using an accessibility overlay widget scored no better than the rest (median 85 against 89).
This one was too low, not too high. We loaded 227 homepages in a real browser and compared them with the crawl's Wappalyzer flags. 23% of the sites flagged as clean did load a tracking script, mostly Google Ads tags and Meta Pixels injected through a tag manager. Detection of consent tools was off in the other direction: 17% of the sites flagged as having none did show a cookie banner.
Lesson: technology detection gives a floor. If a number matters, check a sample by hand and publish the error rate with it.
HTTP Archive doesn't request /llms.txt. We fetched it ourselves: about a third of law firm sites have one. Wix and Duda generate it for nearly every site (99% and 96%), and on WordPress it usually comes from an SEO plugin. Most firms probably don't know the file is there.
The report's charts are plain HTML and CSS with a 2 KB script: no chart library, bars that render without JavaScript, and a tile map for the state data. It seemed wrong to ship a heavy page about heavy pages.
Full numbers, method and limits: The 2026 Law Firm Website Report. Disclosure: Advocentral builds websites for law firms. The report names no individual firm.