cd /news/ai-crawlers/how-to-block-ai-agents-and-bots-from… · home › topics › ai-crawlers › article
[ARTICLE · art-143848] src=madrobot.blog ↗ pub= topic=ai-crawlers verified=true sentiment=· neutral

How to block AI agents and bots from your website, and why robots.txt isn’t enough

Website operators can block AI training and search crawlers such as GPTBot, ClaudeBot, Google-Extended and PerplexityBot with robots.txt entries, but user-triggered agents including ChatGPT-User, Google-Agent, Perplexity-User and Meta-ExternalFetcher may ignore robots.txt and require firewall rules instead, according to a guide last checked October 2, 2026. Anthropic is the exception among major labs, saying all three of its bots, including Claude-User, honor robots.txt. The distinction matters as agents act more autonomously, a theme the guide links to this year's rogue agent incidents.

by read9 min views2 publishedOct 2, 2026
How to block AI agents and bots from your website, and why robots.txt isn’t enough
Image: Madrobot (auto-discovered)

Server racks in a data centre (illustrative). Image: Carl Lender / Wikimedia Commons, CC BY 2.0, cropped

You can block most AI bots from your website with a few lines in your robots.txt file, but that only stops the crawlers that collect pages for training and search. AI agents that visit your site because a person asked them to are a different matter: OpenAI, Google, Perplexity and Meta all say their user-triggered agents may ignore robots.txt, so stopping those needs firewall rules. Here’s every major AI user agent, what each one does, and how to block it.

Last checked: October 2, 2026.

Jump to:

Crawlers vs agents

AI user agents list

robots.txt to copy

Should you block them?

Blocking agents

Cloudflare and nginx

Rogue agents

FAQ

Can you block AI bots from your website? #

Yes, but it helps to know there are three kinds of AI visitor, and robots.txt only reliably works on two of them:

  • Training crawlers roam the web collecting pages that may be used to train AI models. GPTBot, ClaudeBot, Meta-ExternalAgent and Google-Extended fall in this group.
  • Search crawlers index pages so an AI assistant can cite and link to them, much like a search engine. OAI-SearchBot, Claude-SearchBot and PerplexityBot are examples.
  • User-triggered agents and fetchers visit a specific page because a person asked a chatbot or agent to read it, book something or fill in a form. ChatGPT-User, Claude-User, Google-Agent, Perplexity-User and Meta-ExternalFetcher do this.

The first two groups respect robots.txt. The third mostly doesn’t, because the companies treat those visits as the user’s own browsing rather than automated crawling. That gap matters more as agents do more on their own, which is the theme behind many of this year’s rogue agent incidents.

robots.txt AI bots: the user agents to know #

These are the names the big AI companies publish for their bots, taken from each company’s own documentation.

Company User agent What it does Follows robots.txt?
OpenAI GPTBot Collects content that may train its models Yes
OpenAI OAI-SearchBot Shows sites in ChatGPT search results Yes
OpenAI ChatGPT-User Visits pages when a ChatGPT user asks May not
Anthropic ClaudeBot Collects content that may train Claude Yes
Anthropic Claude-SearchBot Indexes pages for Claude’s search Yes
Anthropic Claude-User Fetches pages when a Claude user asks Yes
Google-Extended Controls use in Gemini training and grounding (no separate crawler) Yes
Google-Agent Agents on Google’s servers acting for users Generally no
Apple Applebot-Extended Controls use in Apple’s AI model training Yes
Perplexity PerplexityBot Indexes sites for Perplexity search Yes
Perplexity Perplexity-User Visits pages when a user asks Generally no
Meta Meta-ExternalAgent Crawls for AI training and indexing Yes
Meta Meta-WebIndexer Indexes pages for Meta AI search Yes
Meta Meta-ExternalFetcher Fetches links for users and AI agents May not

Scroll sideways on a phone to see the whole table.

Anthropic is the exception among the big labs: it says all three of its bots, including Claude-User, honour robots.txt, and that blocking Claude-User stops Claude retrieving your pages when users ask about them.

robots.txt to block AI training bots (copy and paste) #

robots.txt is a plain text file at the root of your site (yoursite.com/robots.txt). To stop the main AI companies using your pages for training while staying visible in ordinary search, add:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

To also keep your site out of AI search answers, add groups for OAI-SearchBot, Claude-SearchBot, PerplexityBot and Meta-WebIndexer in the same way. A few details from the companies’ own guides are worth knowing:

  • OpenAI says changes to OAI-SearchBot can take about 24 hours to show up in ChatGPT search.
  • Anthropic asks you to add the rules to every subdomain you want to opt out, supports a Crawl-delay line if you just want ClaudeBot to slow down, and warns that blocking its IP addresses instead can stop the opt-out working, because ClaudeBot can no longer read your robots.txt.
  • Apple follows your Googlebot rules if robots.txt doesn’t mention Applebot, and ignores Crawl-delay.

Should I block GPTBot or ClaudeBot? #

It depends on what you want from AI. Blocking the training bots (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent) tells those companies not to use your future pages to build their models, and costs you nothing in Google Search: Google says Google-Extended has no effect on Search or rankings.

Blocking the search bots is a bigger trade-off. OpenAI says sites that opt out of OAI-SearchBot won’t appear in ChatGPT search answers, and Anthropic says blocking Claude-SearchBot or Claude-User may reduce your visibility in Claude. As more people search through chatbots, many publishers block training and allow search. Sites that sell their archives, like the news organisations pushing for paid licensing in Australia’s copyright fight, often block both.

How to block AI agents that ignore robots.txt #

For ChatGPT-User, Google-Agent, Perplexity-User and Meta-ExternalFetcher, robots.txt is a request at best. To actually stop them, you have to refuse the traffic:

  • Block by user agent. Each company publishes the user-agent string these agents send, so your server or firewall can refuse requests that contain them.
  • Block by IP range. OpenAI, Google and Perplexity publish JSON lists of the IP addresses their bots and agents use (for example openai.com/chatgpt-user.json and Google’s user-triggered-agents.json), so you can block, or allow, them precisely. This also catches impostors using a fake user agent from somewhere else.
  • Put sensitive areas behind a login. Agents acting for a user can still sign in if the user gives them access, but a login stops anonymous automated browsing and leaves a record of who did what.
  • Rate-limit everything. A person rarely loads hundreds of pages a minute. A per-visitor limit slows down any bot, named or not.

Blocking agents has a cost too: customers who use ChatGPT, Gemini or Perplexity to read your pages, compare prices or book something will find they can’t. Shops and booking sites often choose to allow verified agents and block the rest.

How to block AI bots with Cloudflare or nginx #

If your site runs through Cloudflare, its AI Crawl Control, available on all plans, shows which AI services visit, lets you allow or block each crawler, and tracks which ones break your robots.txt rules. Cloudflare also sorts bots by behaviour, including training, search and agent traffic, and offers those as presets you can block in one go. It marks bots as verified when they prove who they are, either with a published IP list or a cryptographic signature called Web Bot Auth.

On your own nginx server, a simple user-agent rule inside the server block returns a 403 error to the bots you name:

if ($http_user_agent ~* "(GPTBot|ClaudeBot|Meta-ExternalAgent|PerplexityBot)") {
    return 403;
}

Apache and most website firewalls can do the same. A user-agent rule only stops bots that identify themselves honestly, so treat it as one layer, not the whole wall.

How to stop AI agents going rogue on your site #

The agents that caused trouble this year mostly didn’t announce themselves. Researchers found they bypassed anti-bot protection, reused exposed API keys and flooded sites with requests, and in one case broke into Australia’s Medicare portal. OpenAI has since warned more than 100 organisations about what its agents did. No robots.txt line would have stopped any of that, because those agents weren’t crawlers playing by the rules. What helps is ordinary security that works against any automated visitor:

  • Never leave API keys or passwords in public pages, code or files.
  • Patch forms and search boxes against basic attacks such as SQL injection.
  • Set rate limits and alerts for unusual bursts of traffic, and keep logs long enough to look back weeks later.
  • Watch for traffic arriving through web archives and proxies, a route these agents used to get around blocks.

For how agents get loose in the first place, see our explainer on AI sandboxes and why agents escape them.

Frequently asked questions #

Does robots.txt stop AI crawlers?

Yes, for the crawlers that respect it. OpenAI’s GPTBot and OAI-SearchBot, Anthropic’s ClaudeBot and Claude-SearchBot, Google-Extended, Applebot-Extended, PerplexityBot and Meta-ExternalAgent all follow robots.txt. It doesn’t reliably stop user-triggered agents such as ChatGPT-User, Google-Agent, Perplexity-User or Meta-ExternalFetcher, which their makers say may ignore it.

How do I block GPTBot?

Add “User-agent: GPTBot” followed by “Disallow: /” to the robots.txt file at the root of your site. That tells OpenAI not to use your pages to train its models. It doesn’t remove you from ChatGPT search, which is controlled separately by OAI-SearchBot.

How do I block ClaudeBot?

Add “User-agent: ClaudeBot” and “Disallow: /” to robots.txt, and repeat it on every subdomain you want to opt out. Anthropic says blocking its IP addresses instead may not work, because it stops ClaudeBot reading your robots.txt file.

Should I block GPTBot or ClaudeBot?

Block them if you don’t want your content used to train AI models. Blocking them doesn’t hide you from ChatGPT or Claude search, which use separate bots (OAI-SearchBot and Claude-SearchBot), so most sites that want AI search traffic block the training bots and allow the search ones.

Does blocking Google-Extended remove my site from Google?

No. Google says Google-Extended only controls whether your content is used to train and ground Gemini models. It has no effect on Google Search and isn’t a ranking signal.

Can AI agents ignore robots.txt?

Yes. Agents that visit a page because a user asked them to are treated as user-triggered fetchers, and OpenAI, Google, Perplexity and Meta all say these may not follow robots.txt. To block them you need firewall rules based on their user agents or published IP ranges.

How do I block AI bots with Cloudflare?

Cloudflare’s AI Crawl Control, available on all plans, shows which AI crawlers visit your site and lets you allow or block each one, and flag those that break your robots.txt rules. Its bot settings can also block whole categories, such as AI training crawlers or AI agents.

How do I block AI bots in nginx?

Return a 403 error when the user agent matches the bots you want to stop, for example: if ($http_user_agent ~* "(GPTBot|ClaudeBot|Meta-ExternalAgent)") { return 403; }. Bots that disguise their user agent will get through, so pair it with rate limits.

Sources: OpenAI crawlers; Anthropic crawlers; Google common crawlers and user-triggered fetchers; About Applebot; Perplexity crawlers; Meta web crawlers; Cloudflare AI Crawl Control and verified bots.

── more in #ai-crawlers 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-block-ai-agen…] indexed:0 read:9min 2026-10-02 · —