# Cloudflare lets publishers block AI training without disappearing from search

> Source: <https://runtimewire.com/article/cloudflare-disallow-ai-training-search-crawler-controls>
> Published: 2026-09-15 13:33:42+00:00

# Cloudflare lets publishers block AI training without disappearing from search

**Cloudflare's new setting publishes a no-training preference while keeping Applebot, Googlebot and Bingbot available for search under its Accountable designation.**

        By [RuntimeWire Staff](https://runtimewire.com/author/runtimewire-staff)
        · Published 

Primary source: [The Cloudflare Blog](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/)

## Why it matters

Mixed-use crawlers have forced publishers to choose between search visibility and control over AI training. Cloudflare's setting reduces that tradeoff by publishing a no-training preference, allowing Accountable crawlers to continue search indexing and blocking other training traffic.

[Cloudflare](https://cloudflare.com/?ref=runtimewire), founded by [Matthew Prince (@eastdakota)](https://x.com/eastdakota?ref=runtimewire), [Michelle Zatlyn (@zatlyn)](https://x.com/zatlyn?ref=runtimewire) and Lee Holloway, launched [Disallow AI Training](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/?ref=runtimewire) on September 15th, 2026, giving website owners a way to reject model training while remaining visible to traditional search crawlers.

The control extends an idea that has run through Cloudflare since its earliest days: the network can enforce a website owner's choices instead of merely recording bad behavior. Prince and Holloway first built Project Honey Pot to track how spammers harvested email addresses. Zatlyn, Prince's Harvard Business School classmate, saw the opening for a service that could stop abuse while improving website performance. Seventeen years later, Cloudflare is applying that same architecture to the increasingly muddled traffic generated by search engines, AI labs and autonomous agents.

Cloudflare's [announcement says fewer than 1% of sites on its network block search bots, while 17% use some mechanism to block AI training](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/?ref=runtimewire). Cloudflare identifies the unit only as "sites" and gives no time period for the measurement, so the percentage is best read as evidence of customer demand rather than a market-wide adoption rate.

### One switch, two enforcement layers

A mixed-use crawler performs search indexing and AI training under one identity. Blocking that crawler at the network level can remove a site from search along with the training pipeline. Allowing it requires the site owner to trust that the operator will observe a preference about how the collected material is used.

Disallow AI Training combines both approaches. Cloudflare's Bot Preference Sync publishes the relevant no-training instruction in `robots.txt`, while Cloudflare's network classifies incoming crawlers and blocks training traffic outside its Accountable category.

Applebot, Googlebot and Bingbot can remain available for search when Cloudflare classifies them as "Accountable" mixed-use crawlers. Training-only crawlers operated by Amazon, Anthropic, Meta and OpenAI are blocked under the setting, according to [Cloudflare](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/?ref=runtimewire); because those crawlers do not also provide search indexing, blocking them does not remove a separate search crawler.

Cloudflare's "Accountable" designation requires an AI-training opt-out, an AI-summary opt-out, URL-level visibility into material made available for training, search-performance reporting and an assurance that rejecting training will not reduce traditional search visibility. The designation recognizes capabilities already available as well as concrete commitments to provide unfinished features. Cloudflare says Apple, Google and Microsoft meet or have committed to meet those qualifications.

### The Accountable badge includes unfinished work

The designation covers capabilities available today alongside promises that Apple, Google and Microsoft have made for future releases. That makes "Accountable" a Cloudflare classification rather than a completed technical standard or independent certification.

[Apple's documentation](https://support.apple.com/en-us/119829?ref=runtimewire) says publishers can disallow `Applebot-Extended` while keeping pages eligible for Applebot search. Neither Apple's documentation nor Cloudflare's announcement establishes a 2027 deadline for additional controls.

[Google's crawler documentation](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers?ref=runtimewire) says `Google-Extended` controls whether crawled material may be used to train future Gemini models and for certain grounding uses. Google says the token does not affect inclusion or ranking in Google Search. The supplied documentation does not establish a deadline for additional URL-level transparency or reporting for appearances in AI summaries.

[Microsoft's Bing guidance](https://blogs.bing.com/webmaster/september-2023/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat?ref=runtimewire) says `NOARCHIVE` excludes material from Bing Chat answers and Microsoft generative-AI model training while allowing it to remain in Bing search results. Neither supplied source establishes an early-2027 deadline for a domain-level Bing preference.

The badge therefore asks site owners to accept Cloudflare's assessment of both present behavior and delivery commitments. Cloudflare says it will publish crawler controls, transparency and reporting through Radar, creating a public record against which those commitments can eventually be measured.

### Cloudflare wants one control panel for the AI web

The September rollout completes the policy-syncing work Cloudflare described when it [introduced Bot Preference Sync in August](https://blog.cloudflare.com/bot-preference-sync/?ref=runtimewire). That earlier feature was designed to keep `robots.txt` declarations aligned with Cloudflare's network enforcement, avoiding cases where a site asked a crawler to stay away while its security configuration continued allowing the traffic.

Cloudflare is now deprecating its older "Block AI Bots" control in favor of separate settings for Search, Training and Agent traffic. Managed Robots.txt is also being replaced by Bot Preference Sync. Existing configurations will generally migrate automatically. For certain new domains, Cloudflare says Disallow AI Training will become part of the recommended configuration.

Cloudflare does not currently offer a Disallow setting for agents because it says the web lacks a well-established directive for expressing that preference. Cloudflare says it may revisit the approach as standards such as [ai-prefs](https://datatracker.ietf.org/wg/aipref/documents/?ref=runtimewire) mature.

Those defaults reveal the commercial theory behind the product. Advertising, subscriptions and direct sales generally require a visitor to reach the website. Model training generates no visit, while an agent can retrieve material without displaying the page to a human. Search remains valuable because it can still deliver the reader or customer.

Cloudflare's position in the request path gives it a practical advantage over policy files alone. A `robots.txt` instruction cannot verify a crawler's identity, determine its purpose or stop an operator that ignores the rule. In its [announcement](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/?ref=runtimewire), Cloudflare says its network identifies who is crawling, classifies why the crawler is visiting and blocks crawlers that ignore the published preference. The system still depends on Cloudflare classifying intent correctly, especially as bots rotate infrastructure or combine multiple purposes under one user agent.

### AI summaries are the next control point

Training preferences settle only one part of the dispute between publishers and AI platforms. Cloudflare plans to address how much content can appear inside AI-generated summaries, where a search or assistant may answer a question before the user visits the source.

By early 2027, Cloudflare aims to give site owners a single setting for controlling the amount of their content used in summaries across participating operators. Cloudflare presents that as a goal, with no finished product or firm launch date in the September 15th announcement.

The distinction matters because publishers may reach different decisions about training, search snippets and generated answers. A retailer may accept a detailed product summary that sends fewer, higher-intent shoppers. An advertising-funded publication may need the page view itself. A domain-wide allow-or-block rule cannot express those differences.

Prince, Zatlyn and Holloway built Cloudflare around the premise that an intermediary network could convert observed abuse into enforceable policy. Cloudflare's crawler controls carry that premise into a web where automated visitors increasingly arrive to index, train, summarize or act. The controls make Cloudflare responsible for classifying those purposes and enforcing each site's chosen setting.
