My AI Crawler Analyzed an Error Page Like It Was the Real Site A developer building a tool that checks how visible webpages are to AI crawlers found it reported a 403 error response from nytimes.com as if it were the real homepage, correctly flagging a missing H1, counting 81 words and finding no JSON-LD in 771 bytes of error content. The fix adds a blocker finding when the HTTP status is 400 or above and stops treating an empty findings list as an all-clear, since the scanner previously never verified that the page it analyzed was the page it requested. I built a small tool recently that checks how visible a webpage is to AI crawlers. You give it a URL, and it tries to fetch the page the way a crawler might. Then it checks things like: noindex ? robots.txt allow GPTBot, ClaudeBot and other AI crawlers? Pretty straightforward. Fetch the page, look at what came back, report what you find. At least that was the idea. Then I pointed it at nytimes.com . My tool reported: Technically, every one of those statements was correct. They were also completely useless. The server had returned HTTP 403 . My tool received 771 bytes of an error response and then happily analyzed it as if it were the New York Times homepage. This is what made the problem interesting to me. The code looking for an H1 worked. There wasn't an H1 in the response, so it correctly reported that there wasn't one. The word counter worked. There were about 81 words, so it reported 81. The structured-data check worked. There wasn't any JSON-LD in the response. Everything did exactly what I had asked it to do. The problem was that I never really asked the most important question: Did I actually get the page I wanted to analyze? The code was basically doing this: js const res = await fetch target, { headers: { 'User-Agent': UA } } ; const html = await res.text ; const page = await extract html ; return report page ; It went straight from "I received something" to "let's analyze it." The funny part is that res.status was already available. I was even displaying it. There was a little grey row near the bottom of the report that said: Status: 403 Meanwhile, above that, the tool was confidently telling me what was supposedly wrong with the page. So technically the information was there. It just wasn't being treated like the most important information on the page. Once I realized what was happening, the first fix was only a few lines: if res.status = 400 { findings.push { level: 'blocker', title: The page returned HTTP ${res.status} to our crawler , detail: 'Everything below was read from the error response, not from the real page.', } ; } There was another problem hiding in the same area. The scanner could also display: Nothing is blocking this page. That message was based on whether the findings list was empty. It wasn't based on whether the page had actually loaded successfully. So under the right conditions, a site could refuse the request completely and my tool could still give it the equivalent of an all-clear. That obviously wasn't what I intended. I'm not a professional developer. I'm mostly building these tools with AI, experimenting, breaking things, fixing them and learning what I should have checked the first time. And I've noticed that a lot of the mistakes I've run into have the same basic shape. I previously had a GPU mining calculator tell me an RTX 4090 could make more than $5,000 a month because I mixed up GH/s and TH/s. I also had calculators where a perfectly valid value of zero was treated as if there was no answer at all. Different bugs, but the same general problem: Something receives an input nobody really questioned, processes it correctly, and produces a very confident answer. Nothing crashes. There isn't necessarily an error message. The page loads. The report looks good. The answer is just wrong. In this case, every part of the scanner worked. It was simply analyzing the wrong thing. The lesson for me wasn't that I need a better HTML parser. It was that before analyzing something, I need to make sure I actually received what I think I received. There are a few surprisingly basic checks that would have caught this. Did the request actually succeed? A 403 response is still a response. JavaScript will happily hand me the body and let me analyze it. Did I get the type of content I expected? If I'm expecting HTML and receive JSON, a PDF, a login page or an error message, searching it for an H1 doesn't tell me much. Does the size make any sense? The response from nytimes.com was 771 bytes. That alone should have made me suspicious that I probably wasn't looking at a major news homepage. Did I end up at the URL I expected? A request can follow redirects and eventually land on a login page, error page or something completely different from the URL that was originally entered. The biggest change I'm making is what happens when one of those checks fails. Instead of producing a normal report with a warning somewhere in it, the tool now needs to say, prominently: I couldn't properly inspect this page, and here's why. Anything it found inside the error response becomes secondary. This part was a little embarrassing. I originally had a much better opening planned. The idea was: nytimes.com loads normally in a browser but returns 403 to an AI crawler. That's a nice example. There's only one problem. Before publishing this, I actually checked it. I requested the same page from the same machine using an ordinary Chrome user agent. 403 again. So I don't actually know that the site was specifically rejecting my crawler. It could be my IP range. It could require more of a normal browser fingerprint. It could be something else entirely. All I can say from my test is: The server refused the requests I made, and my scanner analyzed that refusal as if it were the real webpage. Which is enough to demonstrate the problem. The irony wasn't lost on me. I built a tool because I wanted evidence instead of assumptions, and then nearly opened the article with an assumption I hadn't verified. Checking it took about ninety seconds. One thing I could verify separately was robots.txt . That request returned HTTP 200. Of the ten crawlers my tool checks, eight were disallowed: Googlebot and Bingbot were allowed. That's actually a useful distinction for this kind of tool. A website can make different choices about traditional search engines and AI-related crawlers. Those aren't necessarily the same thing, and putting everything under a single "AI visibility score" can hide that. Which is one reason I don't give the scanner a score out of 100. A page with a slightly short meta description and a page with noindex shouldn't just lose a different number of points. One is a small optimization issue. The other can prevent the page from appearing where you expect it to. The tool is StashGrid AI Visibility https://stashgrid.site/ai-visibility/ . It's free, doesn't require an account and doesn't store the URLs you enter. Paste in a URL and it checks what the crawler actually receives, including: X-Robots-Tag The HTTP response now comes first. If the page returns a 403, that's the headline. Not the missing H1 buried inside the error page. I also rank findings as blocker , warning or info instead of turning everything into one score. And wherever possible, the scanner shows the evidence that caused the finding. I want someone using it to be able to look at the result and decide whether they agree with it. For comparison, I ran it against one of my own pages that it could actually reach. HTTP 200. No redirects. One H1. 2,272 words of readable content. Three JSON-LD blocks. A self-referencing canonical. All ten crawlers allowed. It's a much more boring report. Which is exactly what I want.