Fighting AI slop in science papers ArXiv announced a new rate-limiting policy capping submitters at two submissions per month, after the preprint server received 40,363 submissions in September 2026 — nearly double the 20,569 in September 2024 and more than four times the 9,869 in September 2016. In the announcement blog post, Kat Boboris said the latest month's submissions generated almost 9,000 support tickets and that moderators are seeing a marked increase in dense, AI-written papers, thin papers of narrow scope, and 'salami' papers. The move follows arXiv's May 2026 one-year ban for unchecked AI content and its October 2025 requirement that survey and position papers in the arXiv cs category have peer backing. Anderson's Angle https://www.unite.ai/series/andersons-angle/ Fighting AI Slop in Science Papers Add Unite.AI to your preferred sources on Google https://www.google.com/preferences/source?q=unite.ai AI does everything at scale, from DOS-attack-level web scraping https://www.unite.ai/the-canary-that-reveals-ai-traffic/ :~:text=Sometimes%20constantly to producing, or enabling, such vast volumes https://www.unite.ai/the-survey-paper-ddos-attack-thats-overwhelming-scientific-research/ of scientific research submissions that the result is becoming both a logistical crisis and a crisis of quality https://www.theguardian.com/science/2025/jul/13/quality-of-scientific-papers-questioned-as-academics-overwhelmed-by-the-millions-published , as the signal-to-noise ratio continues to become unnavigable. Last week, preprint server arXiv announced https://blog.arxiv.org/2026/10/01/updated-rate-limit-policy/ that it will be instituting a new rate-limiting policy for submitters, who will now be capped at two submissions per month. The measure has been taken, the announcement suggests, in the face of a vertiginous rise in submission rates over a short time-period: Kat Boboris’s blog post announcing the move, which drew comment https://news.ycombinator.com/item?id=49926512 at Hacker News last Thursday, stated that arXiv received 40,363 submissions in September 2026 – almost double the 20,569 received in September 2024, and more than four times the 9,869 received in September 2016. The latest month’s submissions, Boboris observed, also generated almost 9,000 support tickets for arXiv staff and moderators: ‘Our moderators are observing an increase in thin papers of narrow scope, as well as ‘salami’ papers, where a single work is broken up and submitted as a set of smaller papers. ‘There is also a marked increase in dense, AI-written papers. AI tools are making it easy for authors to flood arXiv and other repositories with these low-value papers.’ Already, in May of this year, arXiv had taken the measure of implementing a one-year ban https://casrai.org/news/arxiv-one-year-ban-unchecked-ai-content for unchecked AI content; and in October of last year, had told submitters of survey papers and position papers both of which can be generated with less effort than a full academic study that they would need peer-backing https://blog.arxiv.org/2025/10/31/attention-authors-updated-practice-for-review-articles-and-position-papers-in-arxiv-cs-category/ in order to be published at arXiv. For the moment, the arXiv domain’s often obstructive capping of http requests is not too much in evidence, and one can only hope that its very useful RSS feeds survive this ongoing retrenchment. Last week Reddit announced https://www.unite.ai/reddit-sets-dates-to-retire-rss-feeds-and-close-public-api-access/ the long-feared total elimination of its RSS feeds, which will take place from the middle of next month. However, arXiv’s non-profit status means that there’s little similar capital https://techcrunch.com/2026/09/30/reddit-is-killing-rss-feeds-ending-public-api-access-because-of-ai-bots/ :~:text=It%E2%80%99s%20no%20wonder%20that%20Reddit%20doesn%E2%80%99t%20want%20to%20give%20any%20of%20that%20data%20away%20for%20free to be gained by shepherding readers into mandatory site visits, or enforced logins https://arstechnica.com/gadgets/2026/06/reddit-will-require-you-to-log-in-to-use-old-reddit-com/ – at least, for the moment. Against the Rising Tide In the face of such severe and growing problems https://www.unite.ai/ai-cites-the-same-papers-over-and-over-again-just-like-humans/ around the negative effect of AI use in science research – most especially regarding AI-related research, which has eclipsed https://cs230.stanford.edu/lecture/1/ :~:text=The,etc%2E%29%2E all other categories, and risen from obscurity to become a political and economic signifier https://www.investing.com/news/commodities-news/trump-signs-order-to-boost-ai-research-using-government-data-4376329 , lately – a strand of research has emerged examining ways to counter the decline in quality of AI-related paper submissions. The latest to address the problem comes in the form of a collaboration between Korea’s Seoul National University and the University of Minnesota in the US. The paper https://arxiv.org/pdf/2610.00531 , titled Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers , proposes measuring ‘scientific slop’ through failures in the connections between a paper’s claims, evidence, citations and structure: Unlike prior approaches, the new system looks beyond the text itself to assess whether the paper’s claims are properly supported by its arguments, evidence and citations. The authors state : ‘Each part of a paper can look plausible in isolation, so such breakdowns are invisible to token-level detectors and can only be identified or repaired at the level of the whole paper. To benchmark and mitigate scientific slop, this paper addresses three challenges. ‘ First, token-level metrics fail to capture how scientific reasoning connects across a paper . Sections, claims, citations, evidence, and artifacts can each appear plausible while the relationships among them break down. We repeatedly observe such failures in end-to-end AI-generated papers, and ICLR reviewers already penalize them even in human-written submissions. ‘ Second, detection of these patterns remains unmeasured . Existing test sets label only the text, so the extent to which detectors, including LLMs that read the entire paper, identify these patterns has never been measured. ‘ Third, these patterns are difficult to mitigate reliably . Whereas token-level signals can be removed by paraphrasing https://arxiv.org/abs/2303.13408 , repairing these patterns requires restoring the missing relations without changing the underlying science.’ The authors’ new benchmark, dubbed SciSlopBench , has been embodied into SciSlopHarness , a framework designed to detect and repair failures in a paper’s scientific reasoning, while preserving the underlying scientific evidence: The SciSlopBench benchmark itself was built from 390 AI-generated papers, each paired with a human-written paper addressing a similar research problem and contribution type. In tests, SciSlopBench was used to distinguish between 390 pairs of papers, with each pair consisting of the aforementioned AI-generated paper, and a human-written paper matched for research problem and type of contribution. The benchmark correctly identified the AI-generated paper in 85.9% of these comparisons, substantially outperforming conventional AI-text detectors. The authors state: ‘While standard revisions leave residual slop and direct slop-aware prompting triggers reward hacking, SciSlopHarness reduces the remaining AI–human gap by 63% over the strongest revision baseline without requiring human reference targets. ‘Overall, we demonstrate that AI-generated scientific papers leave fundamental traces in their global reasoning, and that responsible mitigation demands strict evidentiary grounding rather than mere prose refinement.’ In addition to the contributions made by the paper itself, the authors have operationalized the principles of their new work in the form of a live demo https://yerimoh.github.io/scientific-slop-demo/ where users can submit scientific papers for analysis of the amount of AI slop they contain. Interestingly, the new paper itself, analyzed by the demo, comes in at a ‘moderate’ slop score https://archive.is/ylnvp of 40/100