# Keeping AI search bots in and AI training bots out on a small Next.js site

> Source: <https://dev.to/kacey_3785aafd260f9/keeping-ai-search-bots-in-and-ai-training-bots-out-on-a-small-nextjs-site-4e16>
> Published: 2026-10-08 10:37:40+00:00

I look after the website for [Capital Complete Solutions](https://www.capital-cs.com), a property inventory company in Birmingham where I also work as an inventory clerk. It's a Next.js app on Vercel with Payload CMS, about 30 pages.

This week I wanted to sort out one thing: let AI search tools find and cite our pages, but stop crawlers that only collect training data. I asked on the Cloudflare community and got more useful answers than I expected. Here's what I learned.

The big providers split their crawlers by job, and each one is its own robots.txt token:

| Provider | Training | Search index | Fetch for a user | 
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User | 
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User | 
| Perplexity |  | PerplexityBot | Perplexity-User | 
| Common Crawl | CCBot |  |  | 

If you also want to opt out of Gemini training, Google-Extended is the token for that, and it doesn't affect Google Search. Names do change, so check each provider's own docs before copying this.

Someone on the thread looked at our live robots.txt and pointed out that everything was falling through the `User-agent: *` group, so GPTBot and CCBot were allowed like anyone else.

The fix is named groups. The catch I didn't know: a crawler follows the most specific group that matches its name and ignores `*` completely. So once a bot has its own group, anything you put under `*` no longer applies to it.

In the App Router you can generate robots.txt from `app/robots.ts`:

``` python
import type { MetadataRoute } from 'next'

export default function robots(): MetadataRoute.Robots {
  return {
    rules: [
      { userAgent: ['GPTBot', 'ClaudeBot', 'CCBot'], disallow: '/' },
      { userAgent: '*', allow: '/', disallow: ['/admin'] },
    ],
    sitemap: 'https://www.capital-cs.com/sitemap.xml',
  }
}
```

The search bots aren't named, so they fall into the `*` group and stay allowed.

Well-behaved bots read it. Nothing forces anyone to. If you want enforcement, the traffic has to pass through something that can block it, like Cloudflare's AI Crawl Control. And that only works when Cloudflare is proxying your traffic. DNS only isn't enough, because the requests never touch Cloudflare.

We're on Vercel, which advises against putting another proxy in front of it, so for now robots.txt is what we're relying on.

One reply mentioned another thread where Cloudflare's one-click "Block AI bots" toggle on a Free plan returned a 403 to PerplexityBot as well as the training bots. If being cited matters to you, block bots one at a time instead.

```
curl -s https://www.capital-cs.com/robots.txt
curl -I -A "GPTBot" https://www.capital-cs.com
```

The first shows exactly what bots see. The second only tests user-agent matching, not the bot's real IP. If you're only using robots.txt like us, it still returns a 200 for GPTBot, which is a good reminder of point 3.

Small site, small change, but I'd been assuming "allow all" was the safe default, and for AI training it wasn't what we wanted. If you've split these differently, I'd like to hear how.
