{"slug": "we-audited-8601-law-firm-websites-with-http-archive-three-numbers-were-wrong-we", "title": "We audited 8,601 law firm websites with HTTP Archive. Three numbers were wrong until we checked by hand.", "summary": "Advocentral analyzed 8,601 US law firm homepages from the September 2026 HTTP Archive mobile crawl and found that three crawl-derived figures were wrong until verified by hand. The robots_txt metric's 6.6% AI-crawler Disallow rate overstated blocking — parsing the 716 named files showed only about 1% of sites block an AI crawler site-wide, with most rules coming from stock Squarespace templates — while the 28.5% unlabelled-form-input figure was dropped after Lighthouse's label audit failed on just 6%, and 23% of sites flagged as tracking-free actually loaded Google Ads or Meta Pixel tags. The team also fetched /llms.txt directly, finding roughly a third of firm sites have one, mostly auto-generated by Wix (99%) and Duda (96%).", "body_md": "We wanted to know what the websites of US law firms actually load, so we pulled 8,601 firm homepages out of the public [HTTP Archive](https://httparchive.org/) crawl and analyzed them. The full results are in our [2026 Law Firm Website Report](https://www.advocentral.com/blog/law-firm-website-report-2026). This post is about the method, and about three places where the crawl data gave us a wrong answer.\n\n`httparchive.crawl.pages` (September 2026, mobile, root pages). 9,269 matched; 8,601 were left after removing courts, bar associations, directories and non-US sites.`technologies` (Wappalyzer detections), `custom_metrics.robots_txt`, `custom_metrics.structured_data`, `custom_metrics.a11y` and `summary`.` lighthouse` column across the whole table: the dry run for that one column came back at 17 TiB.\nThe `robots_txt` custom metric counts rules per user agent. 6.6% of sites had a `Disallow` rule for GPTBot, ClaudeBot or another AI crawler, so that looked like the headline.\n\nIt only records that a rule exists, not which path it covers. We fetched and parsed the 716 files that named an AI crawler. About 1% of all sites block one from the whole site. Most of the rest were the stock Squarespace robots.txt, which lists AI crawlers in the same group as `*` and only hides system paths like `/config` and `/search`.\n\n**Lesson:** a rule count is not a block. Parse the file.\n\nThe `a11y` custom metric includes an accessibility tree for form controls. Counting inputs with an empty accessible name gave 28.5% of sites.\n\nThen we ran Lighthouse on 1,954 of the same homepages through the PageSpeed Insights API. Its `label` audit failed on 6%. The likely cause is hidden fields, such as the reCAPTCHA response textarea, being counted as unlabelled inputs. We dropped the 28.5% figure.\n\nWhat Lighthouse did show: 68% of homepages fail color contrast, 54% have links with no accessible name, and the sites using an accessibility overlay widget scored no better than the rest (median 85 against 89).\n\nThis one was too low, not too high. We loaded 227 homepages in a real browser and compared them with the crawl's Wappalyzer flags. 23% of the sites flagged as clean did load a tracking script, mostly Google Ads tags and Meta Pixels injected through a tag manager. Detection of consent tools was off in the other direction: 17% of the sites flagged as having none did show a cookie banner.\n\n**Lesson:** technology detection gives a floor. If a number matters, check a sample by hand and publish the error rate with it.\n\nHTTP Archive doesn't request `/llms.txt`. We fetched it ourselves: about a third of law firm sites have one. Wix and Duda generate it for nearly every site (99% and 96%), and on WordPress it usually comes from an SEO plugin. Most firms probably don't know the file is there.\n\nThe report's charts are plain HTML and CSS with a 2 KB script: no chart library, bars that render without JavaScript, and a tile map for the state data. It seemed wrong to ship a heavy page about heavy pages.\n\nFull numbers, method and limits: [The 2026 Law Firm Website Report](https://www.advocentral.com/blog/law-firm-website-report-2026).\n\n*Disclosure: Advocentral builds websites for law firms. The report names no individual firm.*", "url": "https://wpnews.pro/news/we-audited-8601-law-firm-websites-with-http-archive-three-numbers-were-wrong-we", "canonical_source": "https://dev.to/advocentral/we-audited-8601-law-firm-websites-with-http-archive-three-numbers-were-wrong-until-we-checked-by-34g4", "published_at": "2026-10-06 11:02:33+00:00", "updated_at": "2026-10-06 11:18:12.992906+00:00", "lang": "en", "topics": ["ai-crawlers", "structured-data", "ai-search"], "entities": ["Advocentral", "HTTP Archive", "Wappalyzer", "Lighthouse", "PageSpeed Insights", "GPTBot", "ClaudeBot", "Squarespace"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/we-audited-8601-law-firm-websites-with-http-archive-three-numbers-were-wrong-we", "markdown": "https://wpnews.pro/news/we-audited-8601-law-firm-websites-with-http-archive-three-numbers-were-wrong-we.md", "text": "https://wpnews.pro/news/we-audited-8601-law-firm-websites-with-http-archive-three-numbers-were-wrong-we.txt", "jsonld": "https://wpnews.pro/news/we-audited-8601-law-firm-websites-with-http-archive-three-numbers-were-wrong-we.jsonld"}}