# AI Crawlers Are Hitting Your Site Right Now. Should You Block Them?

> Source: <https://dev.to/searchenginebasics/ai-crawlers-are-hitting-your-site-right-now-should-you-block-them-2e2e>
> Published: 2026-09-30 09:09:52+00:00

Go check your server logs. I'll wait.

See those user agents you don't recognize? `GPTBot`, `ClaudeBot`, `CCBot`, `PerplexityBot`, `Bytespider`, `Amazonbot` — that's not Googlebot doing its usual rounds. That's a second, newer wave of crawlers: bots collecting your content to train AI models and feed retrieval systems that answer questions *without ever sending you a visitor*.

As developers, we've spent a decade optimizing for search crawlers. Now we have to make a call our SEO colleagues can't make for us: **do we let these bots in, or shut the door?** Here's how to think about it like an engineer.

Search crawlers (Googlebot, Bingbot) fetch your pages to build an index that sends you traffic. The deal is obvious: they crawl, you get visitors.

AI crawlers have a murkier value exchange. They generally do one of two things:

The tricky part? The same bot often does both, and the operators are not always transparent about which. That's why a blanket "block everything" or "allow everything" is the wrong move.

This isn't abstract. AI crawlers are aggressive — some ignore `crawl-delay`, hit your site thousands of times a day, and request heavy pages (image galleries, faceted search URLs, API endpoints). I've seen them:

If your site is small, a single aggressive crawler can be a noticeable percentage of your total traffic. Check before you decide anything.

Blocking feels good until you realize what you're giving up. AI-powered answer engines are where a growing share of discovery happens. If your content is never crawled by these systems, it can never be cited by them. For documentation sites, blogs, and open knowledge bases, that's a genuine distribution channel you're closing.

There's also a middle path most teams miss: **you don't have to treat every AI crawler the same.** Retrieval bots that cite sources (and can send traffic) are very different from pure training scrapers. A selective policy beats a binary one.

Here's the framework I use:

For a deeper breakdown of the trade-offs — including which specific bots do what — this guide on [whether you should block AI crawlers](https://www.searchenginebasics.dev/crawling/should-you-block-ai-crawlers/) walks through each major crawler's behavior and the robots.txt directives that control them.

If you decide to block or throttle, do it at the right layer:

**robots.txt** — the standard mechanism. But write it carefully: a sloppy `Disallow` can nuke your search traffic too. Group AI bots explicitly:

```
User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: ClaudeBot
Disallow: /
```

Keep Googlebot and other search crawlers on their own rules, and test the file — one misplaced wildcard has taken down more sites than any algorithm update.

**Server-level controls** — for bots that ignore robots.txt (and some do), enforce at the edge: rate-limit by user agent, return 403s, or challenge with your WAF. Robots.txt is a request, not a lock.

**llms.txt** — the emerging convention: a markdown file at `/llms.txt` that gives AI systems a clean, structured summary of your site. Think of it as a sitemap for language models. Early, but worth watching.

Here's the part most web developers miss: AI crawlers aren't just an SEO problem. The content they collect feeds [enterprise RAG systems](https://www.esaholic.com/services/enterprise-rag-systems/) — retrieval-augmented generation pipelines that companies build on top of web-scale data. Your documentation, your tutorials, your forum answers: they end up chunked, embedded, and retrieved inside someone else's AI product.

That cuts both ways. If you *build* AI products, you should understand exactly how this pipeline consumes web content — because the same crawling, chunking, and embedding mechanics apply to your own knowledge bases. The crawler problem and the RAG problem are the same problem viewed from opposite sides.

The next wave is already here: [AI agents that browse the web like users](https://www.softbrixai.com/services/ai-agent-development/) — clicking, scrolling, filling forms. They don't respect robots.txt the way crawlers do because they often drive real browsers. If you thought bot management was solved, agents are about to reopen the whole discussion: how do you distinguish a helpful agent completing a task for a user from a scraper wearing a browser costume?

For now: get your crawler policy right, instrument your traffic so you can see what's actually hitting you, and keep the policy reviewable. The bot landscape is changing quarterly.

The bots aren't going away. The developers who measure, decide deliberately, and control at the right layer will be fine. The ones running on vibes and blanket blocks will either pay the bandwidth bill or vanish from AI answers entirely. Pick your trade-off on purpose.

*More on crawlers, robots.txt, and technical SEO: browse the [SEO](https://dev.to/t/seo) and [AI](https://dev.to/t/ai) tags here on DEV.*
