cd /news/ai-crawlers/keeping-ai-search-bots-in-and-ai-tra… · home › topics › ai-crawlers › article
[ARTICLE · art-147498] src=dev.to ↗ pub= topic=ai-crawlers verified=true sentiment=· neutral

Keeping AI search bots in and AI training bots out on a small Next.js site

A developer maintaining the Next.js site for Capital Complete Solutions, a Birmingham property inventory company, configured robots.txt named user-agent groups to block AI training crawlers such as GPTBot, ClaudeBot and CCBot while leaving AI search bots like OAI-SearchBot and PerplexityBot allowed under the wildcard group. The writeup notes that a crawler follows only its most specific matching group and ignores the wildcard entirely, and that robots.txt is advisory unless traffic is proxied through a service such as Cloudflare's AI Crawl Control.

by read3 min views3 publishedOct 8, 2026

I look after the website for Capital Complete Solutions, a property inventory company in Birmingham where I also work as an inventory clerk. It's a Next.js app on Vercel with Payload CMS, about 30 pages.

This week I wanted to sort out one thing: let AI search tools find and cite our pages, but stop crawlers that only collect training data. I asked on the Cloudflare community and got more useful answers than I expected. Here's what I learned.

The big providers split their crawlers by job, and each one is its own robots.txt token:

Provider Training Search index Fetch for a user
OpenAI GPTBot OAI-SearchBot ChatGPT-User
Anthropic ClaudeBot Claude-SearchBot Claude-User
Perplexity PerplexityBot Perplexity-User
Common Crawl CCBot

If you also want to opt out of Gemini training, Google-Extended is the token for that, and it doesn't affect Google Search. Names do change, so check each provider's own docs before copying this.

Someone on the thread looked at our live robots.txt and pointed out that everything was falling through the User-agent: * group, so GPTBot and CCBot were allowed like anyone else.

The fix is named groups. The catch I didn't know: a crawler follows the most specific group that matches its name and ignores * completely. So once a bot has its own group, anything you put under * no longer applies to it.

In the App Router you can generate robots.txt from app/robots.ts:

import type { MetadataRoute } from 'next'

export default function robots(): MetadataRoute.Robots {
  return {
    rules: [
      { userAgent: ['GPTBot', 'ClaudeBot', 'CCBot'], disallow: '/' },
      { userAgent: '*', allow: '/', disallow: ['/admin'] },
    ],
    sitemap: 'https://www.capital-cs.com/sitemap.xml',
  }
}

The search bots aren't named, so they fall into the * group and stay allowed.

Well-behaved bots read it. Nothing forces anyone to. If you want enforcement, the traffic has to pass through something that can block it, like Cloudflare's AI Crawl Control. And that only works when Cloudflare is proxying your traffic. DNS only isn't enough, because the requests never touch Cloudflare.

We're on Vercel, which advises against putting another proxy in front of it, so for now robots.txt is what we're relying on.

One reply mentioned another thread where Cloudflare's one-click "Block AI bots" toggle on a Free plan returned a 403 to PerplexityBot as well as the training bots. If being cited matters to you, block bots one at a time instead.

curl -s https://www.capital-cs.com/robots.txt
curl -I -A "GPTBot" https://www.capital-cs.com

The first shows exactly what bots see. The second only tests user-agent matching, not the bot's real IP. If you're only using robots.txt like us, it still returns a 200 for GPTBot, which is a good reminder of point 3.

Small site, small change, but I'd been assuming "allow all" was the safe default, and for AI training it wasn't what we wanted. If you've split these differently, I'd like to hear how.

── more in #ai-crawlers 4 stories · sorted by recency
── more on @capital complete solutions 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/keeping-ai-search-bo…] indexed:0 read:3min 2026-10-08 · —