{"slug": "my-ai-crawler-analyzed-an-error-page-like-it-was-the-real-site", "title": "My AI Crawler Analyzed an Error Page Like It Was the Real Site", "summary": "A developer building a tool that checks how visible webpages are to AI crawlers found it reported a 403 error response from nytimes.com as if it were the real homepage, correctly flagging a missing H1, counting 81 words and finding no JSON-LD in 771 bytes of error content. The fix adds a blocker finding when the HTTP status is 400 or above and stops treating an empty findings list as an all-clear, since the scanner previously never verified that the page it analyzed was the page it requested.", "body_md": "I built a small tool recently that checks how visible a webpage is to AI crawlers.\n\nYou give it a URL, and it tries to fetch the page the way a crawler might. Then it checks things like:\n\n`noindex`?` robots.txt` allow GPTBot, ClaudeBot and other AI crawlers?\nPretty straightforward.\n\nFetch the page, look at what came back, report what you find.\n\nAt least that was the idea.\n\nThen I pointed it at `nytimes.com`.\n\nMy tool reported:\n\nTechnically, every one of those statements was correct.\n\nThey were also completely useless.\n\nThe server had returned **HTTP 403**.\n\nMy tool received 771 bytes of an error response and then happily analyzed it as if it were the New York Times homepage.\n\nThis is what made the problem interesting to me.\n\nThe code looking for an H1 worked.\n\nThere wasn't an H1 in the response, so it correctly reported that there wasn't one.\n\nThe word counter worked.\n\nThere were about 81 words, so it reported 81.\n\nThe structured-data check worked.\n\nThere wasn't any JSON-LD in the response.\n\nEverything did exactly what I had asked it to do.\n\nThe problem was that I never really asked the most important question:\n\n**Did I actually get the page I wanted to analyze?**\n\nThe code was basically doing this:\n\n``` js\nconst res = await fetch(target, { headers: { 'User-Agent': UA } });\nconst html = await res.text();\nconst page = await extract(html);\nreturn report(page);\n```\n\nIt went straight from \"I received something\" to \"let's analyze it.\"\n\nThe funny part is that `res.status` was already available.\n\nI was even displaying it.\n\nThere was a little grey row near the bottom of the report that said:\n\n```\nStatus: 403\n```\n\nMeanwhile, above that, the tool was confidently telling me what was supposedly wrong with the page.\n\nSo technically the information was there.\n\nIt just wasn't being treated like the most important information on the page.\n\nOnce I realized what was happening, the first fix was only a few lines:\n\n```\nif (res.status >= 400) {\n  findings.push({\n    level: 'blocker',\n    title: `The page returned HTTP ${res.status} to our crawler`,\n    detail: 'Everything below was read from the error response, not from the real page.',\n  });\n}\n```\n\nThere was another problem hiding in the same area.\n\nThe scanner could also display:\n\nNothing is blocking this page.\n\nThat message was based on whether the findings list was empty.\n\nIt wasn't based on whether the page had actually loaded successfully.\n\nSo under the right conditions, a site could refuse the request completely and my tool could still give it the equivalent of an all-clear.\n\nThat obviously wasn't what I intended.\n\nI'm not a professional developer.\n\nI'm mostly building these tools with AI, experimenting, breaking things, fixing them and learning what I should have checked the first time.\n\nAnd I've noticed that a lot of the mistakes I've run into have the same basic shape.\n\nI previously had a GPU mining calculator tell me an RTX 4090 could make more than **$5,000 a month** because I mixed up GH/s and TH/s.\n\nI also had calculators where a perfectly valid value of zero was treated as if there was no answer at all.\n\nDifferent bugs, but the same general problem:\n\n**Something receives an input nobody really questioned, processes it correctly, and produces a very confident answer.**\n\nNothing crashes.\n\nThere isn't necessarily an error message.\n\nThe page loads.\n\nThe report looks good.\n\nThe answer is just wrong.\n\nIn this case, every part of the scanner worked.\n\nIt was simply analyzing the wrong thing.\n\nThe lesson for me wasn't that I need a better HTML parser.\n\nIt was that before analyzing something, I need to make sure I actually received what I think I received.\n\nThere are a few surprisingly basic checks that would have caught this.\n\n**Did the request actually succeed?**\n\nA `403` response is still a response. JavaScript will happily hand me the body and let me analyze it.\n\n**Did I get the type of content I expected?**\n\nIf I'm expecting HTML and receive JSON, a PDF, a login page or an error message, searching it for an H1 doesn't tell me much.\n\n**Does the size make any sense?**\n\nThe response from `nytimes.com` was 771 bytes.\n\nThat alone should have made me suspicious that I probably wasn't looking at a major news homepage.\n\n**Did I end up at the URL I expected?**\n\nA request can follow redirects and eventually land on a login page, error page or something completely different from the URL that was originally entered.\n\nThe biggest change I'm making is what happens when one of those checks fails.\n\nInstead of producing a normal report with a warning somewhere in it, the tool now needs to say, prominently:\n\nI couldn't properly inspect this page, and here's why.\n\nAnything it found inside the error response becomes secondary.\n\nThis part was a little embarrassing.\n\nI originally had a much better opening planned.\n\nThe idea was:\n\n`nytimes.com` loads normally in a browser but returns 403 to an AI crawler.\n\nThat's a nice example.\n\nThere's only one problem.\n\nBefore publishing this, I actually checked it.\n\nI requested the same page from the same machine using an ordinary Chrome user agent.\n\n**403 again.**\n\nSo I don't actually know that the site was specifically rejecting my crawler.\n\nIt could be my IP range. It could require more of a normal browser fingerprint. It could be something else entirely.\n\nAll I can say from my test is:\n\n**The server refused the requests I made, and my scanner analyzed that refusal as if it were the real webpage.**\n\nWhich is enough to demonstrate the problem.\n\nThe irony wasn't lost on me.\n\nI built a tool because I wanted evidence instead of assumptions, and then nearly opened the article with an assumption I hadn't verified.\n\nChecking it took about ninety seconds.\n\nOne thing I could verify separately was `robots.txt`.\n\nThat request returned HTTP 200.\n\nOf the ten crawlers my tool checks, eight were disallowed:\n\nGooglebot and Bingbot were allowed.\n\nThat's actually a useful distinction for this kind of tool.\n\nA website can make different choices about traditional search engines and AI-related crawlers. Those aren't necessarily the same thing, and putting everything under a single \"AI visibility score\" can hide that.\n\nWhich is one reason I don't give the scanner a score out of 100.\n\nA page with a slightly short meta description and a page with `noindex` shouldn't just lose a different number of points.\n\nOne is a small optimization issue.\n\nThe other can prevent the page from appearing where you expect it to.\n\nThe tool is **[StashGrid AI Visibility](https://stashgrid.site/ai-visibility/)**.\n\nIt's free, doesn't require an account and doesn't store the URLs you enter.\n\nPaste in a URL and it checks what the crawler actually receives, including:\n\n`X-Robots-Tag`\nThe HTTP response now comes first.\n\nIf the page returns a 403, that's the headline.\n\nNot the missing H1 buried inside the error page.\n\nI also rank findings as **blocker**, **warning** or **info** instead of turning everything into one score.\n\nAnd wherever possible, the scanner shows the evidence that caused the finding.\n\nI want someone using it to be able to look at the result and decide whether they agree with it.\n\nFor comparison, I ran it against one of my own pages that it could actually reach.\n\nHTTP 200.\n\nNo redirects.\n\nOne H1.\n\n2,272 words of readable content.\n\nThree JSON-LD blocks.\n\nA self-referencing canonical.\n\nAll ten crawlers allowed.\n\nIt's a much more boring report.\n\nWhich is exactly what I want.", "url": "https://wpnews.pro/news/my-ai-crawler-analyzed-an-error-page-like-it-was-the-real-site", "canonical_source": "https://dev.to/steven_browning_70ac8fbfa/my-ai-crawler-analyzed-an-error-page-like-it-was-the-real-site-1mdj", "published_at": "2026-10-04 19:37:13+00:00", "updated_at": "2026-10-04 19:42:28.453983+00:00", "lang": "en", "topics": ["ai-crawlers", "ai-tools", "developer-tools"], "entities": ["nytimes.com", "GPTBot", "ClaudeBot"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/my-ai-crawler-analyzed-an-error-page-like-it-was-the-real-site", "markdown": "https://wpnews.pro/news/my-ai-crawler-analyzed-an-error-page-like-it-was-the-real-site.md", "text": "https://wpnews.pro/news/my-ai-crawler-analyzed-an-error-page-like-it-was-the-real-site.txt", "jsonld": "https://wpnews.pro/news/my-ai-crawler-analyzed-an-error-page-like-it-was-the-real-site.jsonld"}}