{"slug": "recursive-generated-entropy", "title": "Recursive Generated Entropy", "summary": "A researcher proposes the term Recursive Generated Entropy (RGE) to describe the progressive loss of independent information, provenance and informational diversity that occurs when generated representations of existing information are recursively consumed as source material for further generation. The paper distinguishes RGE from model collapse, arguing that RGE degrades the information environment itself and can affect capable models, search engines, researchers and human decision-makers. It introduces an independence ratio R = Nₑ/N, illustrating how a single observation reproduced across 999 generated articles yields an apparent 1,000 sources but only one independent lineage.", "body_md": "**Abstract**\n\nGenerative artificial intelligence makes information extraordinarily cheap to reproduce, transform and disseminate. This creates an apparent paradox: the quantity of available information may increase rapidly while the amount of independent information contained within it declines.\n\nThis paper proposes the term **Recursive Generated Entropy (RGE)** for the progressive degradation of an information environment that occurs when generated representations of existing information are repeatedly reused as source material for further generation.\n\nRGE is distinct from model collapse. Model collapse concerns deterioration in models trained recursively on synthetic data. RGE concerns deterioration of the **information environment itself**. It can therefore affect otherwise capable models, search engines, researchers and human decision-makers.\n\nThe central problem is not simply that generated information may be wrong. It is that derivative information can become detached from its provenance and subsequently be mistaken for independent corroboration. Generative systems can therefore produce **reams** of mutually consistent material from remarkably few independent observations.\n\nThe result is an information system in which **volume and apparent certainty increase while provenance, independence and potentially knowledge decline**.\n\n**1. The abundance paradox**\n\nFor most of human history, producing and distributing information was expensive.\n\nWriting a book, publishing a newspaper, maintaining an archive or conducting research required substantial human effort. The cost of reproduction consequently imposed at least some constraint on the quantity of derivative material.\n\nGenerative AI changes that constraint dramatically.\n\nA single source can now produce a summary, which produces an article, which produces hundreds of social-media posts, which are incorporated into databases, which appear in search results, which are summarised by another AI system and ultimately become source material for subsequent generations of models.\n\nThe amount of information apparently available has increased enormously.\n\nBut has the amount of **independent information** increased?\n\nNot necessarily.\n\nConsider an original observation O. Suppose it is copied or transformed into 100 documents.\n\nA search system subsequently encounters 100 documents making substantially the same claim.\n\nThere appear to be 100 sources.\n\nThere may, however, still be only **one observation**.\n\nGenerative AI makes this distinction increasingly important because the cost of creating additional representations is approaching zero.\n\n**2. Recursive Generated Entropy**\n\nI propose the following definition:\n\n**Recursive Generated Entropy (RGE)** is the progressive loss of independent information, provenance and informational diversity that occurs when generated representations of existing information are recursively consumed as source material for further generation.\n\nThe important word is **recursive**.\n\nA conventional copy retains a relatively obvious relationship with its source. Recursive generation creates chains:\n\n**Original → generation → generation of the generation → further generation**\n\nEach transformation may summarise, simplify, reorganise or reinterpret what preceded it.\n\nAfter sufficiently many transformations, the final representation may retain the central claim while losing information about where the claim originated, what qualifications accompanied it and whether apparently separate versions have a common ancestor.\n\nThe genealogy of the information becomes obscure.\n\n**3. Apparent sources and independent sources**\n\nLet **N** be the number of apparent sources and **Nₑ** the number of effectively independent source lineages.\n\nWe can define a simple **independence ratio**:\n\n**R = Nₑ / N**\n\nIf ten genuinely independent witnesses report an event:\n\n**N = 10, Nₑ = 10, R = 1**\n\nNow suppose one report is reproduced by 999 generated articles:\n\n**N = 1,000, Nₑ = 1, R = 0.001**\n\nThe information environment looks dramatically richer.\n\nIts evidential base has not changed.\n\nIndeed, matters may be worse than this simple ratio suggests because successive generations can introduce small mutations. The thousand descendants need no longer make precisely identical claims. Some may acquire additional details through inference, summarisation errors or synthesis with other derivative material.\n\nThose variations can subsequently create the appearance of additional independent evidence.\n\n**The system has begun manufacturing its own corroboration.**\n\n**4. The confidence inversion**\n\nThis produces what may be the most dangerous characteristic of RGE.\n\nOrdinarily, independent corroboration should increase confidence. If genuinely independent sources repeatedly report the same fact, our confidence in that fact should normally rise.\n\nUnder RGE, however, **apparent corroboration can increase as independence decreases**.\n\n**Apparent corroboration ↑ Independent evidence ↓**\n\nA model, researcher or reader unable to reconstruct source genealogy may consequently become **more confident as the quality of the evidential environment deteriorates**.\n\nThis is not principally a hallucination problem.\n\nEvery derivative document might accurately reproduce the original claim.\n\nThe error occurs when **replication is mistaken for corroboration**.\n\n**5. A small accidental experiment**\n\nThe idea for RGE arose from an unexpectedly mundane exercise.\n\nAn attempt was made to reconstruct the career of a former senior employee of an investment-management company approximately twenty years after he had left the industry.\n\nSearches produced several modern financial databases containing records of the individual. The databases appeared superficially to constitute multiple sources.\n\nYet they contained remarkably little information.\n\nThey repeatedly established essentially one fact: the individual had once been associated with the company.\n\nA substantially more informative fact — his actual senior investment role within the organisation — had disappeared from the readily searchable public record.\n\nThe original information environment had therefore undergone an interesting transformation.\n\nThere were **more searchable representations of the surviving fact than there were surviving independent sources**, while important information present in the original environment had vanished.\n\nThe record had become simultaneously **more replicated and less informative**.\n\nGenerative AI did not cause this particular example. Database aggregation and the ordinary decay of the early web were sufficient.\n\nThat is precisely why the example matters.\n\n**Generative AI industrialises the process.**\n\n**6. RGE is not model collapse**\n\nRGE overlaps with, but is distinct from, the phenomenon generally described as **model collapse**.\n\nResearch on recursive synthetic training has demonstrated the possibility that models trained increasingly on model-generated data can lose information about the tails of an original distribution and progressively distort the distribution they attempt to reproduce.\n\nThat is fundamentally a property of a **training process**.\n\nRGE is a property of an **information ecosystem**.\n\nA future model need not have been trained recursively on synthetic data to encounter the problem. Imagine a perfectly capable model with excellent reasoning and retrieval abilities searching an information environment containing ten thousand documents ultimately derived from three primary sources.\n\nUnless it can identify those source relationships, its evidence is compromised.\n\n**Model collapse asks:** what happens when models learn from generated data?\n\n**Recursive Generated Entropy asks:** what happens when civilisation increasingly obtains its information from recursively generated representations of previous information?\n\n**7. Loss of the tails**\n\nSummarisation is necessarily selective.\n\nIf an original document contains 100 facts and a summary retains 30, information has been discarded. A summary of that summary may retain 15. A subsequent synthesis may retain ten.\n\nThe information most likely to survive is generally that which appears central, repeated or conventional.\n\nRare observations, awkward qualifications and apparently peripheral details are disproportionately vulnerable.\n\nConsequently, repeated generation may push an information environment towards its **modal representation**.\n\n**The centre survives. The tails disappear.**\n\nFor knowledge systems, observations in the tails can be exceptionally valuable. They may contain precisely the anomalies that overturn an accepted explanation.\n\n**8. Provenance as information**\n\nTraditional approaches often treat provenance as metadata attached to information.\n\nUnder RGE, provenance should instead be regarded as **part of the information itself**.\n\n“Ten independent primary sources report X”\n\nand\n\n“Ten websites report X”\n\nare informationally different propositions.\n\nIf all ten websites derive from one primary source, collapsing that distinction destroys information relevant to evaluating X.\n\nA robust AI information architecture should therefore preserve not merely documents and citations but **source lineages**.\n\nThe relevant structure is a graph, not a count.\n\nTen nodes with one common parent should not be interpreted in the same way as ten independently rooted nodes.\n\n**9. The economics make recursion inevitable**\n\nRGE is likely to become more important because generative economics strongly favour replication.\n\nThe marginal cost of producing another competent piece of text is approaching zero.\n\nThe marginal cost of producing another **independent observation of reality** is not.\n\nRunning an experiment still costs money. Investigative journalism takes time. Interviewing a witness requires a witness. Examining an archive requires preservation. Measuring the physical world requires instruments. Acquiring genuine expertise may require decades.\n\n**Generation is cheap. Observation remains expensive.**\n\nThere is therefore a powerful economic incentive for the number of representations **N** to grow much faster than the number of independent observations **Nₑ**.\n\nUnless information systems explicitly compensate for this asymmetry, the independence ratio **R** should tend to decline.\n\n**10. Generated Entropy**\n\nRecursive Generated Entropy can therefore produce a peculiar information environment:\n\n**Information volume ↑ Independent information ↓Apparent corroboration ↑Traceability to primary evidence ↓**\n\nThis does not imply that generative AI necessarily destroys knowledge.\n\nAI can equally be used to recover archives, identify common ancestry between sources, compare competing accounts and expose contradictions that humans would struggle to find.\n\nThe technology capable of accelerating RGE may also provide some of the best tools for controlling it.\n\nBut doing so requires recognising the problem.\n\nFuture information systems may need to optimise not simply for the number, relevance or apparent agreement of sources, but for their **effective independence and proximity to primary evidence**.\n\n**11. A recursive experiment**\n\nThere is an obvious experiment embedded in this paper.\n\nAt publication, the term **Recursive Generated Entropy** has a defined meaning and an identifiable primary source.\n\nSuppose the term subsequently propagates.\n\nAI systems may summarise this paper. Those summaries may be published. Other systems may summarise the summaries. Definitions may gradually change. Examples may be added. Attribution may disappear. Multiple websites may eventually present similar definitions without revealing that they share a common origin.\n\nAt some future point, an AI system may be asked:\n\n**Who coined the term Recursive Generated Entropy?**\n\nIts ability — or inability — to reconstruct the answer will itself constitute evidence about the phenomenon described here.\n\nThe concept therefore creates its own longitudinal experiment.\n\n**Conclusion**\n\nGenerative AI promises an unprecedented abundance of information.\n\nAbundance should not be confused with independence.\n\nWhen generated representations recursively become sources for subsequent generation, an information ecosystem can accumulate enormous numbers of mutually reinforcing documents without accumulating corresponding independent evidence.\n\nOver time, provenance may disappear, unusual information may be lost, derivative claims may acquire the appearance of corroboration, and confidence may increase even as the effective evidential base deteriorates.\n\nThat process is **Recursive Generated Entropy**.\n\n**The central challenge for the next generation of information systems may therefore be surprisingly old-fashioned: not producing more information, but remembering where it came from.**", "url": "https://wpnews.pro/news/recursive-generated-entropy", "canonical_source": "https://jakebreams.substack.com/p/recursive-generated-entropy", "published_at": "2026-10-05 22:01:51+00:00", "updated_at": "2026-10-05 22:19:57.183178+00:00", "lang": "en", "topics": ["artificial-intelligence", "generative-ai", "ai-research", "ai-safety", "large-language-models"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/recursive-generated-entropy", "markdown": "https://wpnews.pro/news/recursive-generated-entropy.md", "text": "https://wpnews.pro/news/recursive-generated-entropy.txt", "jsonld": "https://wpnews.pro/news/recursive-generated-entropy.jsonld"}}