AI Is Eating the Web’s Memory — And Making Itself Dumber As of January 2026, 241 news organizations across nine countries, including the New York Times, CNN, USA Today, and The Guardian, have blocked the Internet Archive's Wayback Machine, citing AI training concerns, but the move threatens 2.6 million Wikipedia citations and decades of journalism. The Electronic Frontier Foundation stated in March 2026 that blocking the Archive won't stop AI training but will erase the web's historical record. Meanwhile, Stack Overflow's monthly question volume collapsed 75% from over 200,000 at its 2014 peak to under 50,000 by late 2025, and over 50% of web content is estimated to be AI-generated by 2026, accelerating model collapse. As of January 2026, 241 news organizations across nine countries have blocked the Internet Archive’s Wayback Machine — including the New York Times, CNN, USA Today, and The Guardian. Their stated reason is protecting content from AI training. The actual outcome is that decades of journalism, 2.6 million Wikipedia citations, and much of the web’s institutional memory are quietly vanishing. Publishers trying to protect themselves from AI are accidentally destroying the historical record that archivists, researchers, and developers depend on. Publishers Are Shooting the Wrong Target The core mistake is treating all web crawlers the same. Commercial AI training bots and nonprofit archivists are fundamentally different, but most publisher blocks don’t distinguish between them. When the New York Times implemented what the Wayback Machine’s director called a “hard block” in late 2025, it didn’t meaningfully slow OpenAI’s training operations — OpenAI simply signed licensing deals worth $25 million to $250 million per agreement to secure access to “uncontaminated pre-2022 human data.” The Archive got blocked; the AI companies wrote checks. The Electronic Frontier Foundation made this point bluntly in March 2026: “Blocking the Internet Archive isn’t going to stop AI training. What it will do is ensure that significant chunks of our journalistic record simply disappear.” Wikipedia relies on the Archive for over 2.6 million article citations. Courts use archived pages to settle disputes over what was published and when. Developers use it to find deleted documentation, deprecated APIs, and broken tutorials. Every publisher that blocks the Internet Archive for AI-related reasons https://www.eff.org/deeplinks/2026/03/blocking-internet-archive-wont-stop-ai-it-will-erase-webs-historical-record is destroying something irreplaceable, to stop something they won’t actually stop. AI Training Data Is Eating Itself Stack Overflow’s monthly question volume collapsed 75% — from over 200,000 questions per month at its 2014 peak to under 50,000 by late 2025. Developers stopped asking questions in public because AI tools answer them faster. That sounds like progress. However, the platforms AI was trained on are dying partly because of AI — and the Stack Overflow traffic collapse https://byteiota.com/stack-overflow-traffic-collapses-75-as-ai-replaces-developer-qa/ is one visible symptom of a larger structural problem. By 2026, over 50% of web content is estimated to be AI-generated. The problem: AI models trained on AI-generated content degrade measurably within five training generations. Researchers call this “model collapse.” The ACM was direct: “Model collapse is already happening. We just pretend it isn’t.” https://cacm.acm.org/blogcacm/model-collapse-is-already-happening-we-just-pretend-it-isnt/ A Nature study confirmed the mechanism — each generation of AI-on-AI training amplifies errors and eliminates the rare-but-important patterns that make expert knowledge valuable. The ouroboros is real: AI trains on content produced by AI trained on content AI consumed, and the loop tightens with each iteration. The companies accelerating this crisis understand it. AI labs are racing to license “uncontaminated pre-2022 human data” at prices ranging from $25 million to $250 million per deal. They’re paying a premium for authentic human knowledge while their crawlers simultaneously reduce the incentive to produce more of it. When AI gives free answers, fewer people write detailed technical posts. When fewer posts get written, the next model’s training data gets thinner. The economics are self-defeating, and no one is stopping. The Scale of Link Rot and Web Memory Loss Link rot is not a new problem, but it is an accelerating one. Old Dominion University analyzed 27.3 million URLs archived since 1996 and found 65% are now dead. Pew Research found https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears/ that 25% of all web pages from 2013 to 2023 are no longer accessible — and 54% of Wikipedia’s citations point to pages that no longer exist. Permanent link rot tripled from 5% in 2012 to 15% in 2025. The web’s knowledge graph is rotting faster than it is growing. For developers, the consequence is immediate. AI tools trained on the degrading web produce answers that cite sources that no longer exist or have moved behind paywalls. Documentation disappears. API references go dead. Stack traces reference libraries at URLs that 404. The Wayback Machine was the fallback — the place to find what was. Fewer publishers are allowing that fallback to exist, as documented by the Nieman Journalism Lab in January 2026 https://www.niemanlab.org/2026/01/news-publishers-limit-internet-archive-access-due-to-ai-scraping-concerns/ . What Developers Can Do About It The open web does not maintain itself. Developers built much of it through public Stack Overflow answers, open GitHub repos, public documentation, and freely indexed blog posts. The trend away from public contribution is not inevitable — it is a choice shaped by incentives that can shift. Publishing technical knowledge openly rather than behind corporate portals, contributing to open documentation projects, and donating to the Internet Archive are concrete actions that counter the trend directly. The EU AI Act, enforced as of August 2026, now requires AI developers to disclose training data sources and respect copyright opt-outs — creating regulatory levers that didn’t exist two years ago. The Harvard Journal of Law and Technology has argued for a legal right to uncontaminated human-generated training data. These are early signals that the industry recognizes the problem, even if the current response is uncoordinated and often counterproductive. The direction of change is not locked in. The web’s memory is worth defending. Key Takeaways - Publishers blocking the Internet Archive to stop AI training are destroying decades of historical record without meaningfully slowing AI companies, which simply buy licensed data instead - AI trained on AI-generated content degrades within five generations — model collapse is measurable and already documented; the web’s quality crisis is structural, not cyclical - 65% of URLs archived since 1996 are dead; 25% of web pages from 2013-2023 are gone; the web’s memory is rotting faster than it grows - Stack Overflow’s 75% question volume collapse shows how AI consumption of human knowledge reduces the incentive to produce new human knowledge — a feedback loop with no natural floor - Developers have agency: publishing openly, supporting open archives, and using EU AI Act transparency rights can slow the trend without waiting for structural fixes