{"slug": "when-search-eats-the-web-a-model-of-corpus-erosion-under-generative-extraction", "title": "When Search Eats the Web: A Model of Corpus Erosion Under Generative Extraction", "summary": "A new arXiv paper models the web corpus as a common-pool resource and proves that generative search engines degrade its volume, quality, and lifetime, potentially causing extinction. The authors show that a myopic engine can cross the erosion threshold while a long-run oriented one stays below, and that adding competing engines increases extraction rates. They also prove the socially optimal extraction rate lies below the threshold and discuss seven survival mechanisms.", "body_md": "# Computer Science > Computer Science and Game Theory\n\n[Submitted on 16 Aug 2026]\n\n# Title:When Search Eats the Web: A Model of Corpus Erosion under Generative Extraction\n\n[View PDF](/pdf/2608.15896)\n\n[HTML (experimental)](https://arxiv.org/html/2608.15896v1)\n\nAbstract:Generative search engines (GSEs) answer user queries directly from crawled web content. The capture of value from the corpus without a visit returned to the source (we call this capture extraction) diverts the traffic that finances content production. In response, publishers may restrict crawler access to their websites. In this paper, we model the crawlable corpus as a common-pool resource: the crawlable commons. It is described by three quantities: volume, average quality, and lifetime. Under two types of responses of publishers we prove that extraction degrades all three at once: publishers opt out, renewal loses its funding, and content becomes more perishable. After a given erosion threshold, the corpus goes extinct. A myopic GSE can cross this threshold, a long-run oriented GSE stays below it. We extend our model to several competing engines and prove, under a concavity condition on the steady-state value of the commons, that the symmetric equilibrium extraction rate is nondecreasing in their number and converges to the threshold. Adding users who strictly prefer direct answers, the assumption most favorable to extraction, we prove that the socially optimal extraction rate lies strictly below the erosion threshold, and no higher than the single engine's sustainable optimum. Finally, we discuss seven survival mechanisms.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/when-search-eats-the-web-a-model-of-corpus-erosion-under-generative-extraction", "canonical_source": "https://arxiv.org/abs/2608.15896", "published_at": "2026-08-18 07:40:52+00:00", "updated_at": "2026-08-18 08:11:13.407461+00:00", "lang": "en", "topics": ["artificial-intelligence", "generative-ai", "ai-policy"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/when-search-eats-the-web-a-model-of-corpus-erosion-under-generative-extraction", "markdown": "https://wpnews.pro/news/when-search-eats-the-web-a-model-of-corpus-erosion-under-generative-extraction.md", "text": "https://wpnews.pro/news/when-search-eats-the-web-a-model-of-corpus-erosion-under-generative-extraction.txt", "jsonld": "https://wpnews.pro/news/when-search-eats-the-web-a-model-of-corpus-erosion-under-generative-extraction.jsonld"}}