cd /news/artificial-intelligence/recursive-generated-entropy · home › topics › artificial-intelligence › article
[ARTICLE · art-145700] src=jakebreams.substack.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Recursive Generated Entropy

A researcher proposes the term Recursive Generated Entropy (RGE) to describe the progressive loss of independent information, provenance and informational diversity that occurs when generated representations of existing information are recursively consumed as source material for further generation. The paper distinguishes RGE from model collapse, arguing that RGE degrades the information environment itself and can affect capable models, search engines, researchers and human decision-makers. It introduces an independence ratio R = Nₑ/N, illustrating how a single observation reproduced across 999 generated articles yields an apparent 1,000 sources but only one independent lineage.

read9 min views1 publishedOct 5, 2026
Recursive Generated Entropy
Image: source

Abstract

Generative artificial intelligence makes information extraordinarily cheap to reproduce, transform and disseminate. This creates an apparent paradox: the quantity of available information may increase rapidly while the amount of independent information contained within it declines.

This paper proposes the term Recursive Generated Entropy (RGE) for the progressive degradation of an information environment that occurs when generated representations of existing information are repeatedly reused as source material for further generation.

RGE is distinct from model collapse. Model collapse concerns deterioration in models trained recursively on synthetic data. RGE concerns deterioration of the information environment itself. It can therefore affect otherwise capable models, search engines, researchers and human decision-makers.

The central problem is not simply that generated information may be wrong. It is that derivative information can become detached from its provenance and subsequently be mistaken for independent corroboration. Generative systems can therefore produce reams of mutually consistent material from remarkably few independent observations.

The result is an information system in which volume and apparent certainty increase while provenance, independence and potentially knowledge decline.

1. The abundance paradox

For most of human history, producing and distributing information was expensive. Writing a book, publishing a newspaper, maintaining an archive or conducting research required substantial human effort. The cost of reproduction consequently imposed at least some constraint on the quantity of derivative material.

Generative AI changes that constraint dramatically.

A single source can now produce a summary, which produces an article, which produces hundreds of social-media posts, which are incorporated into databases, which appear in search results, which are summarised by another AI system and ultimately become source material for subsequent generations of models.

The amount of information apparently available has increased enormously.

But has the amount of independent information increased?

Not necessarily.

Consider an original observation O. Suppose it is copied or transformed into 100 documents.

A search system subsequently encounters 100 documents making substantially the same claim.

There appear to be 100 sources.

There may, however, still be only one observation.

Generative AI makes this distinction increasingly important because the cost of creating additional representations is approaching zero.

2. Recursive Generated Entropy

I propose the following definition:

Recursive Generated Entropy (RGE) is the progressive loss of independent information, provenance and informational diversity that occurs when generated representations of existing information are recursively consumed as source material for further generation.

The important word is recursive.

A conventional copy retains a relatively obvious relationship with its source. Recursive generation creates chains:

Original → generation → generation of the generation → further generation

Each transformation may summarise, simplify, reorganise or reinterpret what preceded it.

After sufficiently many transformations, the final representation may retain the central claim while losing information about where the claim originated, what qualifications accompanied it and whether apparently separate versions have a common ancestor.

The genealogy of the information becomes obscure.

3. Apparent sources and independent sources

Let N be the number of apparent sources and Nₑ the number of effectively independent source lineages. We can define a simple independence ratio:

R = Nₑ / N

If ten genuinely independent witnesses report an event:

**N = 10, Nₑ = 10, R = 1**

Now suppose one report is reproduced by 999 generated articles:

N = 1,000, Nₑ = 1, R = 0.001 The information environment looks dramatically richer.

Its evidential base has not changed.

Indeed, matters may be worse than this simple ratio suggests because successive generations can introduce small mutations. The thousand descendants need no longer make precisely identical claims. Some may acquire additional details through inference, summarisation errors or synthesis with other derivative material.

Those variations can subsequently create the appearance of additional independent evidence.

The system has begun manufacturing its own corroboration.

4. The confidence inversion

This produces what may be the most dangerous characteristic of RGE.

Ordinarily, independent corroboration should increase confidence. If genuinely independent sources repeatedly report the same fact, our confidence in that fact should normally rise.

Under RGE, however, apparent corroboration can increase as independence decreases.

Apparent corroboration ↑ Independent evidence ↓

A model, researcher or reader unable to reconstruct source genealogy may consequently become more confident as the quality of the evidential environment deteriorates.

This is not principally a hallucination problem.

Every derivative document might accurately reproduce the original claim.

The error occurs when replication is mistaken for corroboration.

5. A small accidental experiment

The idea for RGE arose from an unexpectedly mundane exercise.

An attempt was made to reconstruct the career of a former senior employee of an investment-management company approximately twenty years after he had left the industry.

Searches produced several modern financial databases containing records of the individual. The databases appeared superficially to constitute multiple sources.

Yet they contained remarkably little information.

They repeatedly established essentially one fact: the individual had once been associated with the company.

A substantially more informative fact — his actual senior investment role within the organisation — had disappeared from the readily searchable public record.

The original information environment had therefore undergone an interesting transformation.

There were more searchable representations of the surviving fact than there were surviving independent sources, while important information present in the original environment had vanished.

The record had become simultaneously more replicated and less informative.

Generative AI did not cause this particular example. Database aggregation and the ordinary decay of the early web were sufficient.

That is precisely why the example matters.

Generative AI industrialises the process.

6. RGE is not model collapse

RGE overlaps with, but is distinct from, the phenomenon generally described as model collapse.

Research on recursive synthetic training has demonstrated the possibility that models trained increasingly on model-generated data can lose information about the tails of an original distribution and progressively distort the distribution they attempt to reproduce.

That is fundamentally a property of a training process.

RGE is a property of an information ecosystem.

A future model need not have been trained recursively on synthetic data to encounter the problem. Imagine a perfectly capable model with excellent reasoning and retrieval abilities searching an information environment containing ten thousand documents ultimately derived from three primary sources.

Unless it can identify those source relationships, its evidence is compromised.

Model collapse asks: what happens when models learn from generated data?

Recursive Generated Entropy asks: what happens when civilisation increasingly obtains its information from recursively generated representations of previous information?

7. Loss of the tails

Summarisation is necessarily selective.

If an original document contains 100 facts and a summary retains 30, information has been discarded. A summary of that summary may retain 15. A subsequent synthesis may retain ten. The information most likely to survive is generally that which appears central, repeated or conventional.

Rare observations, awkward qualifications and apparently peripheral details are disproportionately vulnerable.

Consequently, repeated generation may push an information environment towards its modal representation.

The centre survives. The tails disappear.

For knowledge systems, observations in the tails can be exceptionally valuable. They may contain precisely the anomalies that overturn an accepted explanation. 8. Provenance as information

Traditional approaches often treat provenance as metadata attached to information.

Under RGE, provenance should instead be regarded as part of the information itself.

“Ten independent primary sources report X”

and

“Ten websites report X”

are informationally different propositions.

If all ten websites derive from one primary source, collapsing that distinction destroys information relevant to evaluating X. A robust AI information architecture should therefore preserve not merely documents and citations but source lineages.

The relevant structure is a graph, not a count.

Ten nodes with one common parent should not be interpreted in the same way as ten independently rooted nodes.

9. The economics make recursion inevitable

RGE is likely to become more important because generative economics strongly favour replication.

The marginal cost of producing another competent piece of text is approaching zero.

The marginal cost of producing another independent observation of reality is not.

Running an experiment still costs money. Investigative journalism takes time. Interviewing a witness requires a witness. Examining an archive requires preservation. Measuring the physical world requires instruments. Acquiring genuine expertise may require decades.

Generation is cheap. Observation remains expensive.

There is therefore a powerful economic incentive for the number of representations N to grow much faster than the number of independent observations Nₑ.

Unless information systems explicitly compensate for this asymmetry, the independence ratio R should tend to decline.

10. Generated Entropy

Recursive Generated Entropy can therefore produce a peculiar information environment:

Information volume ↑ Independent information ↓Apparent corroboration ↑Traceability to primary evidence ↓

This does not imply that generative AI necessarily destroys knowledge.

AI can equally be used to recover archives, identify common ancestry between sources, compare competing accounts and expose contradictions that humans would struggle to find.

The technology capable of accelerating RGE may also provide some of the best tools for controlling it.

But doing so requires recognising the problem.

Future information systems may need to optimise not simply for the number, relevance or apparent agreement of sources, but for their effective independence and proximity to primary evidence.

11. A recursive experiment

There is an obvious experiment embedded in this paper.

At publication, the term Recursive Generated Entropy has a defined meaning and an identifiable primary source.

Suppose the term subsequently propagates.

AI systems may summarise this paper. Those summaries may be published. Other systems may summarise the summaries. Definitions may gradually change. Examples may be added. Attribution may disappear. Multiple websites may eventually present similar definitions without revealing that they share a common origin.

At some future point, an AI system may be asked:

Who coined the term Recursive Generated Entropy?

Its ability — or inability — to reconstruct the answer will itself constitute evidence about the phenomenon described here.

The concept therefore creates its own longitudinal experiment.

Conclusion

Generative AI promises an unprecedented abundance of information.

Abundance should not be confused with independence.

When generated representations recursively become sources for subsequent generation, an information ecosystem can accumulate enormous numbers of mutually reinforcing documents without accumulating corresponding independent evidence.

Over time, provenance may disappear, unusual information may be lost, derivative claims may acquire the appearance of corroboration, and confidence may increase even as the effective evidential base deteriorates.

That process is Recursive Generated Entropy.

The central challenge for the next generation of information systems may therefore be surprisingly old-fashioned: not producing more information, but remembering where it came from.

── more in #artificial-intelligence 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/recursive-generated-…] indexed:0 read:9min 2026-10-05 · —