{"slug": "keeping-ai-search-bots-in-and-ai-training-bots-out-on-a-small-next-js-site", "title": "Keeping AI search bots in and AI training bots out on a small Next.js site", "summary": "A developer maintaining the Next.js site for Capital Complete Solutions, a Birmingham property inventory company, configured robots.txt named user-agent groups to block AI training crawlers such as GPTBot, ClaudeBot and CCBot while leaving AI search bots like OAI-SearchBot and PerplexityBot allowed under the wildcard group. The writeup notes that a crawler follows only its most specific matching group and ignores the wildcard entirely, and that robots.txt is advisory unless traffic is proxied through a service such as Cloudflare's AI Crawl Control.", "body_md": "I look after the website for [Capital Complete Solutions](https://www.capital-cs.com), a property inventory company in Birmingham where I also work as an inventory clerk. It's a Next.js app on Vercel with Payload CMS, about 30 pages.\n\nThis week I wanted to sort out one thing: let AI search tools find and cite our pages, but stop crawlers that only collect training data. I asked on the Cloudflare community and got more useful answers than I expected. Here's what I learned.\n\nThe big providers split their crawlers by job, and each one is its own robots.txt token:\n\n| Provider | Training | Search index | Fetch for a user | \n|---|---|---|---|\n| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User | \n| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User | \n| Perplexity |  | PerplexityBot | Perplexity-User | \n| Common Crawl | CCBot |  |  | \n\nIf you also want to opt out of Gemini training, Google-Extended is the token for that, and it doesn't affect Google Search. Names do change, so check each provider's own docs before copying this.\n\nSomeone on the thread looked at our live robots.txt and pointed out that everything was falling through the `User-agent: *` group, so GPTBot and CCBot were allowed like anyone else.\n\nThe fix is named groups. The catch I didn't know: a crawler follows the most specific group that matches its name and ignores `*` completely. So once a bot has its own group, anything you put under `*` no longer applies to it.\n\nIn the App Router you can generate robots.txt from `app/robots.ts`:\n\n``` python\nimport type { MetadataRoute } from 'next'\n\nexport default function robots(): MetadataRoute.Robots {\n  return {\n    rules: [\n      { userAgent: ['GPTBot', 'ClaudeBot', 'CCBot'], disallow: '/' },\n      { userAgent: '*', allow: '/', disallow: ['/admin'] },\n    ],\n    sitemap: 'https://www.capital-cs.com/sitemap.xml',\n  }\n}\n```\n\nThe search bots aren't named, so they fall into the `*` group and stay allowed.\n\nWell-behaved bots read it. Nothing forces anyone to. If you want enforcement, the traffic has to pass through something that can block it, like Cloudflare's AI Crawl Control. And that only works when Cloudflare is proxying your traffic. DNS only isn't enough, because the requests never touch Cloudflare.\n\nWe're on Vercel, which advises against putting another proxy in front of it, so for now robots.txt is what we're relying on.\n\nOne reply mentioned another thread where Cloudflare's one-click \"Block AI bots\" toggle on a Free plan returned a 403 to PerplexityBot as well as the training bots. If being cited matters to you, block bots one at a time instead.\n\n```\ncurl -s https://www.capital-cs.com/robots.txt\ncurl -I -A \"GPTBot\" https://www.capital-cs.com\n```\n\nThe first shows exactly what bots see. The second only tests user-agent matching, not the bot's real IP. If you're only using robots.txt like us, it still returns a 200 for GPTBot, which is a good reminder of point 3.\n\nSmall site, small change, but I'd been assuming \"allow all\" was the safe default, and for AI training it wasn't what we wanted. If you've split these differently, I'd like to hear how.", "url": "https://wpnews.pro/news/keeping-ai-search-bots-in-and-ai-training-bots-out-on-a-small-next-js-site", "canonical_source": "https://dev.to/kacey_3785aafd260f9/keeping-ai-search-bots-in-and-ai-training-bots-out-on-a-small-nextjs-site-4e16", "published_at": "2026-10-08 10:37:40+00:00", "updated_at": "2026-10-08 10:49:10.396825+00:00", "lang": "en", "topics": ["ai-crawlers", "generative-engine-optimization", "ai-search", "developer-tools"], "entities": ["Capital Complete Solutions", "Next.js", "Vercel", "Payload CMS", "Cloudflare", "GPTBot", "ClaudeBot", "CCBot"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/keeping-ai-search-bots-in-and-ai-training-bots-out-on-a-small-next-js-site", "markdown": "https://wpnews.pro/news/keeping-ai-search-bots-in-and-ai-training-bots-out-on-a-small-next-js-site.md", "text": "https://wpnews.pro/news/keeping-ai-search-bots-in-and-ai-training-bots-out-on-a-small-next-js-site.txt", "jsonld": "https://wpnews.pro/news/keeping-ai-search-bots-in-and-ai-training-bots-out-on-a-small-next-js-site.jsonld"}}