Major publishers are increasingly limiting OpenAI's GPTBot from accessing their reporting, marking a broader shift in how news organizations assert control over content used for AI training. The BBC and The Guardian list GPTBot as disallowed in their robots.txt policies, while The New York Times has also prohibited scraping for AI training and development without explicit permission in its terms of service.
The development matters because web crawling has long been a route to assembling large training datasets. When high-profile publishers restrict access at the source, AI developers face a more constrained and more clearly governed data environment. The issue is not simply whether a crawler can retrieve a page. It is increasingly about permission, licensing and accountable data provenance.
The Guardian's published robots.txt directives provide a direct example of this approach. The file disallows GPTBot alongside a broader set of bots, signaling that the publisher does not want its content scraped for AI training or data aggregation.
Robots.txt is a machine-readable file that tells web crawlers which parts of a site they are permitted to access. For AI-related crawlers, it has become a practical opt-out mechanism. Publishers are pairing that technical control with contractual restrictions and discussions around licensing, rather than relying on informal expectations about how online content may be reused.
The actions documented across major publishers are not identical, but they point in the same direction: indiscriminate collection of publisher content is becoming harder to justify and operationalize. The distinction is important because some publisher policies differentiate between crawlers used for model training and systems used for retrieval, indexing or other purposes.
| Publisher | Documented action | Relevant implication |
|---|---|---|
| The Guardian | Its robots.txt disallows GPTBot and a broader set of bots. | Signals restrictions on AI training or data-aggregation scraping. |
| BBC | Its robots.txt lists GPTBot and other AI-related crawlers as disallowed. | Formally limits AI data collection from BBC properties. |
| The New York Times | In August 2023, it updated its terms of service to prohibit scraping for AI training and development without explicit permission. | Adds a contractual restriction alongside publicly observed crawler blocks. |
Reuters Institute research published across 2023 and 2024 identified a cluster of leading publishers, including the BBC, The New York Times, CNN and Reuters, that had begun blocking GPTBot access. The Guardian also reported that outlets including ABC and the Chicago Tribune were taking similar steps. Together, these actions indicate an industry-level consolidation of data-rights controls, rather than an isolated policy decision by one publisher.
For OpenAI and other model developers, GPTBot restrictions narrow one potential path for collecting web material. GPTBot is OpenAI's crawler for gathering training data for models such as ChatGPT. The practical result is not that all public web content becomes unavailable to AI systems. Instead, it increases the importance of distinguishing permitted sources from restricted sources and of documenting how data was acquired. The publisher response brings three operational questions into sharper focus:
This shift also affects enterprises that are not building foundation models. Organizations developing internal AI tools may use external datasets, third-party models or retrieval systems whose data practices have different constraints. Procurement, legal, security and technical teams therefore need a shared view of what content enters an AI workflow, under which terms, and for what purpose.
The central uncertainty is how publisher restrictions, opt-out mechanisms and prospective licensing arrangements will align with evolving AI-rights frameworks and enforcement practices. The available evidence does not establish a single industry standard. It does show that major publishers are making their preferences more explicit through both technical and contractual mechanisms.
Organizations assessing AI data sources, retrieval architectures or vendor controls can work with Scalevise on AI governance, workflow design and implementation decisions that account for data provenance and permission boundaries.
What is GPTBot?
GPTBot is OpenAI's web crawler for gathering data used to train models such as ChatGPT.
Which publishers have blocked GPTBot?
The verified material identifies the BBC and The Guardian as listing GPTBot as disallowed in robots.txt. Reuters Institute research also noted blocks by leading publishers including The New York Times, CNN and Reuters.
Does The New York Times prohibit AI training on its content?
The New York Times updated its terms of service in August 2023 to prohibit scraping its content for AI training and development without explicit permission.
Do robots.txt restrictions apply to every AI use case?
Not necessarily. Publisher policies can distinguish training crawlers from retrieval, indexing or other bots, so each AI use case requires separate review.
Why do GPTBot blocks matter to enterprise AI teams?
They make data provenance, permissions and licensing more important when teams select datasets, build retrieval systems or assess AI vendors.
The BBC, The Guardian and The New York Times illustrate a wider publisher effort to control how journalism is used in AI development. As GPTBot restrictions and related terms become more common, model builders and enterprise teams will need more rigorous approaches to permissions, licensing and data governance.