{"slug": "robots-txt-for-ai-crawlers-gptbot-claudebot-perplexitybot-explained", "title": "robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot explained", "summary": "A developer published a guide explaining how to use robots.txt to control AI crawlers, distinguishing training crawlers like GPTBot, ClaudeBot and Google-Extended from search and answer crawlers such as OAI-SearchBot and PerplexityBot. The guide provides copy-paste rules for three policies — allow everything, block training while allowing search, and block all listed AI crawlers — and notes that robots.txt is a request rather than access control, since scrapers that ignore it are not stopped. It also includes a Python snippet using urllib.robotparser to verify rules and warns that its wildcard handling differs from Google's.", "body_md": "AI companies now run their own crawlers, and most identify themselves with a user-agent name you can put in robots.txt. This guide explains what the main ones do, how training crawlers differ from search and answer crawlers, and gives copy-paste rules for three common policies. Confirm details against each vendor's current documentation before you deploy.\n\nFirst, a limit: robots.txt is a request, not access control. Well-behaved crawlers honor it. Scrapers that ignore it are not stopped by it, and a blocked URL can still appear in search results if other sites link to it.\n\n**GPTBot** is OpenAI's crawler. OpenAI documents it as collecting web content that may be used to train its models.\n\n**OAI-SearchBot** is OpenAI's crawler for ChatGPT search. It indexes pages so ChatGPT can show them in search results and cite them. It is separate from GPTBot, so you can allow one and block the other.\n\n**ChatGPT-User** runs when a person asks ChatGPT to open a page during a conversation. It is user-triggered rather than a bulk crawl, and OpenAI notes that robots.txt rules may not apply to these requests.\n\n**ClaudeBot** is Anthropic's crawler, used to collect content that may be used for training. Anthropic also documents separate agents for search and user-requested fetching, so check its current list before writing rules for them.\n\n**PerplexityBot** indexes pages for Perplexity's answers and citations. Perplexity says it respects robots.txt. Its user-triggered fetcher, Perplexity-User, behaves like ChatGPT-User.\n\n**Google-Extended** is not a crawler. It is a robots.txt token that controls whether content Google has already crawled can be used to train Gemini and for certain grounding features. Googlebot still crawls your site for Search whether or not you block this token, so blocking Google-Extended does not remove you from Google Search.\n\nThe useful split is purpose.\n\nIf you want citations and referral traffic from AI answers, allow the search crawlers. If your objection is only to training, block the training agents and keep the search agents open.\n\n**Allow everything.** This is the default, written out explicitly:\n\n```\nUser-agent: *\nAllow: /\n```\n\n**Block training, allow search and answers:**\n\n```\nUser-agent: GPTBot\nUser-agent: ClaudeBot\nUser-agent: Google-Extended\nDisallow: /\n\nUser-agent: OAI-SearchBot\nUser-agent: PerplexityBot\nAllow: /\n\nUser-agent: *\nAllow: /\n```\n\n**Block all of these AI crawlers:**\n\n```\nUser-agent: GPTBot\nUser-agent: OAI-SearchBot\nUser-agent: ChatGPT-User\nUser-agent: ClaudeBot\nUser-agent: PerplexityBot\nUser-agent: Google-Extended\nDisallow: /\n\nUser-agent: *\nAllow: /\n```\n\nA crawler obeys the group that names it and ignores the `*` group. So the bots you block stay blocked even though `*` allows everything. The `*` group applies to crawlers you did not name, such as Googlebot and Bingbot.\n\n`User-agent: *` and `Disallow: /`.` https://example.com/robots.txt`. Each subdomain, such as `blog.example.com`, needs its own file.`/Blog/` and `/blog/` are different paths. Bot names are case-insensitive, but copy the vendor's spelling anyway.\n\n``` python\nfrom urllib.robotparser import RobotFileParser\n\nrp = RobotFileParser(\"https://example.com/robots.txt\")\nrp.read()\nfor bot in [\"GPTBot\", \"OAI-SearchBot\", \"PerplexityBot\"]:\n    print(bot, rp.can_fetch(bot, \"https://example.com/blog/post\"))\n```\n\nIts wildcard handling differs from Google's, so verify key paths with a second parser.\n\nDecide your policy first: allow all, block training only, or block everything. Write named groups for the bots you care about, host the file at the root, and verify it with logs and a parser.\n\nIf you would rather generate these groups than type them, the free generator at [https://citescore.vercel.app/robots-txt-ai-crawler-generator](https://citescore.vercel.app/robots-txt-ai-crawler-generator) builds them.", "url": "https://wpnews.pro/news/robots-txt-for-ai-crawlers-gptbot-claudebot-perplexitybot-explained", "canonical_source": "https://dev.to/citescore/robotstxt-for-ai-crawlers-gptbot-claudebot-perplexitybot-explained-46mc", "published_at": "2026-10-10 20:43:44+00:00", "updated_at": "2026-10-10 20:46:13.855272+00:00", "lang": "en", "topics": ["ai-crawlers", "generative-engine-optimization", "ai-search", "ai-tools"], "entities": ["OpenAI", "GPTBot", "ClaudeBot", "Anthropic", "PerplexityBot", "Perplexity", "Google-Extended", "Google"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/robots-txt-for-ai-crawlers-gptbot-claudebot-perplexitybot-explained", "markdown": "https://wpnews.pro/news/robots-txt-for-ai-crawlers-gptbot-claudebot-perplexitybot-explained.md", "text": "https://wpnews.pro/news/robots-txt-for-ai-crawlers-gptbot-claudebot-perplexitybot-explained.txt", "jsonld": "https://wpnews.pro/news/robots-txt-for-ai-crawlers-gptbot-claudebot-perplexitybot-explained.jsonld"}}