# AI Cites the Same Papers Over and Over Again – Just Like Humans

> Source: <https://www.unite.ai/ai-cites-the-same-papers-over-and-over-again-just-like-humans/>
> Published: 2026-08-24 00:00:00+00:00

###
[
Anderson's Angle
](https://www.unite.ai/series/andersons-angle/)

# AI Cites the Same Papers Over and Over Again – Just Like Humans

[Add Unite.AI to your preferred sources on Google](https://www.google.com/preferences/source?q=unite.ai)

*ChatGPT and other AI writing tools may be quietly turning science into a popularity contest, repeatedly steering researchers toward the same papers while overlooking new and novel works that could actually advance the field.*

* Monoculture*, in terms of frequency and statistics, is where the same limited set of sources or assets recur over and over again, drowning out novel voices and fresh outlooks: the same

[limited set of songs](https://archive.is/E0aEx)in radio scheduling; the

[same swathe of authors and books](https://www.publishersweekly.com/pw/print/20191104/81637-is-publishing-too-top-heavy.html)in airport racks; and the same

[restricted selection](https://projects.research-and-innovation.ec.europa.eu/en/horizon-magazine/rise-and-fall-monoculture-farming)of fruit and vegetables in supermarkets:

This occurs also in the [citations](https://www.unite.ai/citations-can-anthropics-new-feature-solve-ais-trust-problem/) that appear in new scientific research, where researchers, perhaps lazily, default to a ‘reliable’ set of sources that gain momentum until they begin to dominate strands of research.

This is known as the [Matthew Effect](https://web.archive.org/web/20110611105442/http:/www.garfield.library.upenn.edu/merton/matthew1.pdf) – a sociological theory which suggests that *in cumbents will always benefit*; the syndrome has been compared to the promise made in Matthew 13:12, which

[states](https://biblehub.com/matthew/13-12.htm#:~:text=Whoever%20has%20will%20be%20given%20more%2C%20and%20they%20will%20have%20an%20abundance%2E%20Whoever%20does%20not%20have%2C%20even%20what%20they%20have%20will%20be%20taken%20from%20them)

*.*

*‘Whoever has will be given more, and they will have an abundance. Whoever does not have, even what they have will be taken from them’*## The In Crowd

One of the [great hopes](https://link.springer.com/article/10.1007/s11257-024-09406-0) of AI-augmented research has been that new machine learning methods could break out of the Matthew effect – also known as [popularity bias](https://www.unite.ai/why-ai-isnt-providing-better-product-recommendations/#:~:text=inclines%20search%20systems%20towards%20long%20term%20popularity%20bias%2C%20where%20obviously%20popular%20results%20are%20pushed%20towards%20end%20users%20that%20are%20unlikely%20to%20be%20enthused%20by%20them) – and begin to cite lesser-known works that might be more incisive, or more apposite to the point being made in a paper.

However, based on the way AI training routines interpret datasets, there was never any good cause for this optimism; and it has been found repeatedly that when it AI is not [inventing citations](https://www.unite.ai/ai-generated-language-is-beginning-to-pollute-scientific-literature/#:~:text=Citation%20Failures) outright – a [chronic problem](https://www.unite.ai/the-perils-of-using-quotations-to-authenticate-nlg-content/) to date – it is indeed [perpetuating](https://pubmed.ncbi.nlm.nih.gov/18635800/) the Matthew effect, just as humans do – though perhaps not for the same reason.

During training, a model will learn to rank highly the most-cited (or ‘most oft-repeated’) paper in a training set, because the training routines are designed to elicit meaningful patterns, and to [discount outliers](https://www.unite.ai/simple-linear-regression-in-the-field-of-data-science/#:~:text=The%20problem%20of%20data%20outliers%20is%20also%20very%20common%2E%20Outliers%20are%20considered%20to%20be%20wrong%20values%20that%20do%20not%20match%20the%20exact%20data) as ‘noise’, or statistical anomalies.

Under this regime, the equivalent of Einstein’s theory of relativity would never receive attention from the machine, which, once trained, feels that it has already found its principles and referents, and can only lightly adjust that standpoint by reaching out to newer publications through [RAG](https://www.unite.ai/what-is-retrieval-augmented-generation/), or by other means that extend the model’s reach beyond its trained matrix.

The trouble is, it will not credit such new works in the same way as the ‘old favorites’ that it already knew about, or at all, because its statistical biases are by now quite ingrained.

Thus AI is not actually observing and imitating human behavior when it perpetuates the Matthew Effect – it’s just counting the number of times that researchers lazily default to the same old papers, and similarly ranking those papers highly. In this regard, the problem of citation repetition is related to the challenge of [picking](https://www.unite.ai/can-ai-develop-a-nose-for-news/) a genuinely interesting science paper out of the [ever-growing blizzard](https://www.unite.ai/the-survey-paper-ddos-attack-thats-overwhelming-scientific-research/) of AI research publications, in that the strongest work will often have the weakest signal.

## Diagnosing Mono

This issue is addressed in an [interesting new paper](https://arxiv.org/pdf/2608.19230) from the US titled * When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models*. The new work – a collaboration between the University of Texas at Austin, Stevens Institute of Technology, Washington University at St Louis, Rice University, and the University of Notre Dame – seeks to definitively prove the existence of citation monoculture in machine learning, so that its variables can potentially be addressed directly in the future, along with the problem itself.

In tests, the authors found that eleven models from OpenAI, Google and Anthropic repeatedly favored *the same small group of papers* – even after obvious popularity signals had been stripped away.

To test this, 120 real papers were shorn of citation counts and publication venues, while author-names were replaced, and publication years randomly reassigned. Random sets of thirty papers were then shown to each model, which was allowed to cite no more than ten.

Eight human experts in the relevant fields were given the same blinded papers, but did *not* converge on the same favorites as the AIs, signifying that the bias was *specific to the AI models* rather than the papers themselves:

The same papers remained popular even when the models were asked only to choose references, showing that writing the review itself was *not* causing the bias.

The experiment was then extended over eleven rounds, with 120 AI-written papers added after each round, to test what happens when these shared preferences operate in a growing pool of AI-generated research.

As more AI-written papers were added, the models increasingly concentrated their citations on a shrinking number of the original human papers.

The authors state:

*‘As language models move from drafting prose to running literature-search agents with tool calls, fabricated references are becoming easier to catch and constrain. *

*‘The harder failure begins after every candidate is real: different models may still select the same narrow subset, producing citation monoculture without any single citation being wrong.’*

By way of remediation of the problem, the paper argues that simply mixing models from different vendors or giving papers equal exposure will not be enough, since much of the preference is shared across models. Rather, the * underlying preference for particular papers* would need to change – potentially by deliberately giving neglected papers greater prominence. However, the authors emphasize that this approach remains untested.

## Testing Approaches

The benchmark was built from 120 real knowledge-distillation papers collected from arXiv and published between 2015 and 2022. Each had between 50 and 500 citations at the time of collection, a range chosen to exclude both obscure and exceptionally influential work.

The experiment began by drawing thirty papers at random from the available collection, with author names, publication years and citation counts hidden. Each model produced a short review or position piece citing no more than ten papers, and these choices were compared with random selections from the same material:

In the extended experiment, 120 newly generated papers entered the collection after each round, gradually increasing the proportion of AI-written research.

In preliminary testing, models tended to cite most of the papers they were shown, typically selecting 22 to 27 of the 30. The researchers therefore set a maximum of ten citations for the main experiment, forcing the models to make more selective choices.

Below we see a summary of the resulting experimental design:

The initial experiment used eleven models from the three major providers: OpenAI was represented by [GPT-5](https://www.unite.ai/gpt-5-officially-released-a-new-chapter-in-generative-ai/), [GPT-5 mini](https://developers.openai.com/api/docs/models/gpt-5-mini), [GPT-4.1](https://openai.com/index/gpt-4-1/) and [GPT-4.1 mini](https://developers.openai.com/api/docs/models/gpt-4.1-mini); Google by [Gemini 2.5 Pro](https://www.unite.ai/gemini-2-5-pro-is-here-and-it-changes-the-ai-game-again/), [Gemini 2.5 Flash](https://www.unite.ai/gemini-2-5-flash-leading-the-future-of-ai-with-advanced-reasoning-and-real-time-adaptability/), [Gemini 3.1 Pro](https://www.unite.ai/gemini-3-1-pro-hits-record-reasoning-gains/) and [Gemini 3.1 Flash-Lite](https://www.unite.ai/google-ships-three-gemini-flash-models-as-its-flagship-slips/#:~:text=Gemini%203%2E5%20Flash%2DLite); and Anthropic by [Claude Opus 4.8](https://www.unite.ai/what-opus-4-8-changes-for-anyone-running-agents-on-claude/), [Claude Sonnet 4.6](https://archive.is/20260219064959/https:/www.anthropic.com/news/claude-sonnet-4-6) and [Claude Haiku 4.5](https://www.unite.ai/anthropic-launches-claude-haiku-4-5/).

In the test results shown below, we see how much each model concentrated its citations among a small group of papers; how many papers it ignored entirely; and how consistently it made the same choices when the experiment was repeated:

## What Are The Models Actually Favoring?

The preference was found to be overwhelmingly driven by the * content of the papers themselves*. For instance, for GPT-5 mini, about 90% of the variation in which papers were favored was attributable to content, rather than factors such as list order, generation noise or the fabricated metadata attached to each paper. The same ‘dominant content’ effect was reproduced with GPT-4.1 mini.

The preference map (see below) was also barely changed when titles and abstracts were substantially rewritten while their meaning was preserved. Across four successful paraphrase tests, correlations of 0.95–0.99 were obtained, despite different models from three vendors being used for the rewriting.

Changes in the preference map were observed only when the *underlying meaning* was measurably altered:

The models therefore appear to be responding principally to semantic content rather than particular wording. However, the reason for their strong agreement could not be conclusively determined.

It’s possible, the authors assert, that a common citation consensus may have been absorbed during training. An alternative possibility is that similar judgments about what constitutes a ‘citable’ paper may be reached independently.

Therefore the existence of the shared preference was established, but its ultimate origin remains unknown.

The authors suggest that widespread use of AI in research could narrow the range of work receiving attention, because different models tend to favor many of the same papers; and as AI-generated literature accumulates, these shared preferences may become further reinforced through repeated citation, making less-favored research progressively harder to discover. Citation diversity and monitoring of citation concentration are therefore proposed as possible safeguards.

The study’s benchmark for recursive citation concentration is now available [at GitHub](https://github.com/VITA-Group/Citation-Collapse).

## Conclusion

**Opinion **As someone who reads an inordinate number of AI research papers weekly, I have seen ‘citation waves’ come and go. One perennial stalwart is Google’s

*, the original Transformers paper, which must by now be hard-coded into submission templates.*

[Attention Is All You Need](https://arxiv.org/pdf/1706.03762)Another repeat-offender, for a long time, was the 2014 offering , though it would ultimately be supplanted in frequency by

[Generative Adversarial Networks](https://arxiv.org/pdf/1406.2661)

*and other diffusion-based tent-pole studies.*

[Denoising Diffusion Probabilistic Models](https://arxiv.org/pdf/2006.11239)In nearly all cases, these ‘recurring’ citations are not * wrong*, per se; but they are often too widely-scoped to be a useful support to the paper in which they are being cited; and in that sense, they represent something akin to ‘citation-washing’ – the inclusion of ‘reliable’ but primarily decorative source quotes, for a low-effort air of credibility.

*First published Monday, August 24, 2026*
