cd /news/ai-policy/cloudflare-ai-crawler-controls-three… · home topics ai-policy article
[ARTICLE · art-74430] src=byteiota.com ↗ pub= topic=ai-policy verified=true sentiment=· neutral

Cloudflare AI Crawler Controls: Three Switches, One Deadline

Cloudflare launched granular AI traffic controls this month, replacing its single 'block AI bots' toggle with three independently configurable categories—Search, Agent, and Training—available to all customers including Free plan users. AI crawlers now account for 57.5% of HTML web traffic across Cloudflare's network, with training bots returning almost no value: ClaudeBot crawls 23,951 pages for every single referral it sends back, while Google's search crawler crawls 4.9 pages per referral. Starting September 15, 2026, new domains will have Training and Agent bots blocked by default on ad-supported pages, though existing domains are grandfathered.

read4 min views1 publishedJul 26, 2026
Cloudflare AI Crawler Controls: Three Switches, One Deadline
Image: Byteiota (auto-discovered)

Cloudflare launched granular AI traffic controls this month, replacing its single “block AI bots” toggle with three independently configurable categories: Search, Agent, and Training — available to all customers, including Free plan users. The announcement landed with a number that explains everything: AI crawlers now account for 57.5% of HTML web traffic across Cloudflare’s network, and training bots return almost nothing in exchange. Cloudflare’s own crawl-to-referral data shows ClaudeBot crawls 23,951 pages for every single referral it sends back. Google’s traditional search crawler, for comparison, crawls 4.9 pages per referral. Website owners have been subsidizing AI training at extraordinary rates, and now Cloudflare is handing them a switch.

Search, Agent, Training: What Each Cloudflare Switch Controls #

The three categories map to how AI companies actually use your content. Search covers bots that index your content so AI tools can answer questions about it — in theory, you get referral traffic or citations in return. Agent covers real-time automated systems acting on users’ behalf: ChatGPT’s browsing mode, browser-use agents, agentic workflows that fetch live web data. Training covers crawlers permanently absorbing your content into model weights — the category with the worst return-on-access ratio by a wide margin.

For each category, site owners get three policy options: allow on all pages, block on all pages, or block only on ad-supported pages. The last option reflects Cloudflare’s stated logic: “An ad is a signal that a website owner meant for a person to land there and see it.” A developer running documentation can allow Agent bots for user convenience while blocking Training. A news publisher can allow Search and block Training. Before this, you could block everything or block nothing.

September 15, 2026: The Date That Changes Cloudflare Defaults #

Starting September 15, 2026, every new domain added to Cloudflare will have Training and Agent bots blocked by default on ad-supported pages. Search stays allowed. Existing domains are grandfathered — nothing changes unless you opt in. Developers launching new sites after that date should understand the defaults before they go live.

There is a real complication worth knowing: multi-purpose crawlers. Googlebot crawls for both traditional search indexing and for training Gemini. Block “Training” and you may inadvertently limit Googlebot’s indexing activity too, since Cloudflare cannot yet separate the two uses from a single crawler. For sites heavily dependent on Google search traffic, that tension matters. The controls give you a choice — but the choice is not perfectly clean when crawlers serve multiple masters.

ClaudeBot Crawls 23,951 Pages Per Referral — Here Is Why That Number Matters #

The crawl-to-referral numbers from Cloudflare’s Radar analysis make the policy inevitable. Of all AI crawler requests across Cloudflare’s network, 51.8% are for Training purposes. Only 9.3% are for Search. The bots consuming the most bandwidth return the least value. ClaudeBot’s 23,951:1 ratio is the outlier, but Perplexity’s 111:1 is still a far cry from Google’s 4.9:1. These are not edge cases — they represent the dominant pattern of how AI companies currently access web content.

Cloudflare CEO Matthew Prince noted the 57.5% bot traffic milestone “arrived years before he expected” — he had forecast this threshold for late 2027. Moreover, agentic AI traffic grew 7,851% year-over-year according to separate analysis. The economics of the open web were heading toward a breaking point regardless of whether Cloudflare acted.

The Bigger Story: Cloudflare Is Building a Toll Booth for AI #

The three-category switch is the immediate news. Pay Per Crawl is the structural shift beneath it. Cloudflare launched a private beta monetization layer where, instead of blocking AI crawlers, publishers can charge them. AI companies register with Ed25519 cryptographic signatures, declare a maximum price they’ll pay via a crawler-max-price

header, and receive a 402 response if their price falls short. Cloudflare acts as Merchant of Record, aggregating billing and distributing earnings to publishers. Current beta partners include Ceramic.ai and You.com, per Cloudflare’s Pay Per Crawl announcement.

A parallel “Pay Per Use” model — charging when content creates value in an AI response, not just when fetched — is also under development. Cloudflare is in discussions with news organizations, publishers, and large-scale social media platforms. The company controlling CDN access for roughly 20% of all websites is positioning itself as the economic gatekeeper between AI companies and the content they need. Whether that is a feature or a concern depends entirely on which side of the transaction you sit on.

Key Takeaways #

  • Cloudflare replaced its blunt AI bot toggle with three independently configurable categories — Search, Agent, and Training — available on all plans including Free
  • Starting September 15, 2026, new Cloudflare domains will have Training and Agent bots blocked by default on ad-supported pages; existing domains are unaffected
  • Training bots represent 51.8% of AI crawler traffic but return almost no referral value — ClaudeBot crawls nearly 24,000 pages per referral sent
  • Multi-purpose crawlers like Googlebot create a real complication: blocking Training may affect standard search indexing until Cloudflare can separate the two uses
  • Pay Per Crawl (private beta) lets publishers charge AI companies for access rather than block them — Cloudflare is positioning itself as an economic intermediary between AI companies and web content
── more in #ai-policy 4 stories · sorted by recency
── more on @cloudflare 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cloudflare-ai-crawle…] indexed:0 read:4min 2026-07-26 ·