List of companies that block Internet Archive An incomplete list compiled by Consumer Rights Wiki identifies companies that have blocked the Internet Archive's Wayback Machine, with 23 major news outlets including The New York Times and USA Today updating their robots.txt files in April 2023 to explicitly block the crawler, according to a report by Am John Philip. The blocking is largely motivated by publishers' desire to prevent their content from being used to train AI models. Inherently incomplete. Sourced to https://archive.is/Lz4nY | ← Older revision | Revision as of 21:27, 17 August 2026 | || | Line 1: | Line 1: | || {{Incomplete}} | {{Incomplete}} | || This is | This is an early/inherently incomplete list of companies who took the lead their websites purposefully blacklisted/excluded from the Internet Archive IA for various reasons, whether for malice reasons plausible deniability or not. For these websites, utilize other archive websites found Consumer Rights Wiki:Tools for writing articles Archived citations|here . | || Unless there's a way for IA to have the content but not make it available to train AI, a large proportion of content sites will be blocking it, in order to avoid their content being available to train AI. So the list is very incomplete. For example, it was reported in April of 2023 "Twenty-three major news outlets, including The New York Times and USA Today have updated their robots.txt files to explicitly block the Internet Archive’s Wayback Machine crawler."