AI-Written Web Text Helps Small Models, Then Starts to Hurt A paper submitted to arXiv on September 30, 2026 by researchers at the University of Maryland and AI detection company Pangram Labs found that 31.1% of quality-filtered August 2026 web tokens are AI-generated, up from 10.1% two years earlier, and that pretraining 800 small models on mixtures of human and AI text shows AI text lowers loss only for models trained on fewer than about 10 human tokens per parameter. At the compute-optimal budget of 20 human tokens per parameter the benefit disappears, and beyond it AI text raises loss, with the paper's scaling law putting training on unfiltered web at today's AI share at 1.6 times the compute of training on its human subset. Repeating human text for eight added epochs lowered loss by 6.3% to 6.4% versus adding the same volume of AI text. A paper submitted to arXiv on September 30, 2026 asks a question every lab now faces: what happens when the web you crawl for training data is partly written by the models you trained last year? The authors, from the University of Maryland and the AI detection company Pangram Labs, label almost a third of quality-filtered August 2026 web tokens as AI-generated, then pretrain 800 small models on mixtures of human and AI text. Their finding is a sign change. AI text helps a model that lacks human data, stops helping at the compute-optimal budget, and raises loss beyond it. Editorial note: Prepared October 3 as an October 1, 2026 dispatch from the paper’s arXiv abstract and HTML text, version 1. One paper, not yet peer reviewed. It makes no claim about how search engines treat AI text, and neither does this post. 1. 0131.1% of filtered web tokensThe share of tokens passing FineWeb’s quality filter that Pangram labels AI-generated, for August 2026. Two years earlier it was 10.1%. 2. 02Help below 10 tokens per parameterAI text lowers loss on human text only for models trained on fewer than about 10 human tokens per parameter. At the standard 20, the benefit is gone. 3. 031.6 times the computeBy the paper’s law, training on unfiltered web at today’s AI share costs 1.6 times the compute of training on its human subset, and the gap grows. 4. 04Repeating human text beats adding AI textAfter eight added epochs, loss was 6.3 to 6.4% lower when human data was repeated than when the same volume of AI text was added. 01 — The measurementHow much of the web is AI-written now The paper, How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text https://arxiv.org/abs/2609.40295 , starts by measuring. The authors take Common Crawl snapshots, apply the FineWeb quality filter that many open pretraining pipelines use, and run two detectors over what survives: a small open classifier to label everything, then Pangram’s own detector as a second check on the confident cases. The result is the chart below. The two later bars are the authors’ forecasts, not measurements. Share of quality-filtered web tokens labelled AI-generated arXiv:2609.40295 v1, figure 1 and section 5.1. Measured shares are Pangram labels on FineWeb-filtered Common Crawl; the 2027 and 2028 figures are the paper’s forecasts. Two things about the measurement matter for everything after it. The share is of tokens that pass a quality filter, and the paper reports that AI text survives filtering more often than human text does, so the filtered share is higher than the raw share. And the label comes from a detector sold by one of the authors’ employers. The paper is open about both; a reader should carry them forward. 02 — The experimentWhat 800 models showed The experiment holds a human corpus fixed and adds AI text to it, up to 64 AI tokens for every human token, across models from 19.9 million to 973 million parameters. Loss is measured on held-out human text from three sources and, separately, on held-out AI text. The design is what distinguishes the paper from earlier model-collapse work: the AI text is not generated by the model under test or curated to help, it is whatever the web contained, from many models, written for people. 19.9M to 973M parameters Varying human tokens per parameter and the ratio of AI to human tokens. Wild AI, released with the paper 42.2B human and 35.3B AI tokens, labelled by source, topic and format. Below this, AI text helps At the compute-optimal 20, the benefit is gone; above it, AI tokens raise loss. The pattern the authors describe is consistent across sizes. For a model starved of human data, adding AI text lowers loss on human text at first, then the benefit saturates and reverses as more is added. For a model already trained at or past the compute-optimal budget of about 20 human tokens per parameter, AI text raises loss almost immediately, and does so more sharply at larger sizes, while the same number of fresh human tokens keeps lowering it. The paper notes that today’s production models train far past that budget, citing one open model at 1,100 tokens per parameter, which is the regime where the harm applies. 03 — The modelA scaling law where a token can hurt Existing scaling laws cannot express this, and the paper says why. The Chinchilla law treats an AI token as a human token. Laws built for repeated data let a token’s value fall towards zero but never below it. The authors fit a law with separate benefit and harm terms, so the value of an AI token can change sign, and which reduces to Chinchilla when no AI text is present. Fit on the smaller models, it predicts the effect of AI text on models up to 3.6 times larger with 41% lower error than the best existing law, on the authors’ own evaluation. The law yields the paper’s most quotable figure. Training on unfiltered web text at August 2026’s AI share requires 1.6 times the compute of training on the human subset alone, at the compute-optimal budget, and the gap widens with more human data. At the forecast shares of 42% and 51%, the multipliers become 2.1 and 3.0. Those last two numbers rest on a forecast and a fitted law, and should be quoted as such. 04 — RecommendationsWhat the authors recommend The abstract ends with four recommendations for anyone building a pretraining corpus. They are short enough to table. | Source: arXiv:2609.40295 v1, abstract and section 5, September 30, 2026. Paraphrased; the supporting figures are the paper’s. | | | |---|---|---| | Recommendation | When | The paper’s support | |---|---|---| | Filter AI text out | When the target is human-written text. | The 1.6 times compute gap at today’s share, growing with the human budget. | | Repeat human text before adding AI text | When the human corpus runs out before the compute does. | Loss 2.3% lower after one added epoch of repeated human data than after the same volume of AI data, widening to 6.3 to 6.4% after eight. | | Report validation loss on human and AI text separately | Always. | At a 22.3% AI share in the evaluation crawl, a mixed validation set hid the harm in 95.5% of the runs where harm occurred. | | Keep AI text when the target is AI text | When the model will mostly read or score machine output. | AI tokens keep lowering loss on held-out AI text; the two are treated as separate domains. | The third row is the one with consequences outside research labs. If a mixed validation set hides the harm, then a team that fine-tunes on scraped web text and checks only an aggregate loss will not see the problem it has introduced. Our decision guide on synthetic training data https://www.digitalapplied.com/blog/synthetic-data-generation-llm-training-decision-guide-2026 makes the same point from the curated side: generated data is a tool with a target, not a free supply. 05 — FencesWhat the study cannot say It measures loss, not capability. Every result is a held-out cross-entropy on text; the paper does not report whether the models answer questions worse or follow instructions worse. Lower loss on human text is the standard proxy, and the authors use it as one, but a reader should not translate the 1.6 figure into a benchmark score. Its largest model is under one billion parameters. The scaling law is validated on models 3.6 times larger than those it was fit on, which is still far below production scale. Whether the harm term keeps its shape at tens of billions of parameters is a prediction the paper makes, not a measurement it reports. Its labels come from a detector. A token is “AI-generated” when Pangram says so, and the company that sells Pangram co-authored the paper. The authors disclose this and release the corpus with its labels so others can relabel it. Until someone does, the share figures are one detector’s view. And it says nothing about search. The paper is about what AI text does to a model that trains on it. It does not test, and does not claim, that search engines rank AI-written pages differently; the question of what publishers should do with AI text is covered in our note on Google’s updated guidance on reviewing AI content https://www.digitalapplied.com/blog/google-ai-content-fact-check-guidance-2026 , which is a separate matter with a separate primary. For the data-supply side, our piece on publishers and Common Crawl https://www.digitalapplied.com/blog/publishers-common-crawl-ai-training-data-showdown-2026 is the context for who decides what gets crawled. The practical lesson transfers down from pretraining. If your fine-tuning set includes scraped web pages, forum threads or documentation written since 2024, assume a material share is model-written, hold out human-written and machine-written examples separately, and look at both losses. The paper’s evidence says the aggregate number can look fine while the half you care about gets worse. Our AI transformation https://www.digitalapplied.com/services/ai-transformation work builds that split into the evaluation from the start. Split your validation set before you trust the loss Whatever you train, keep a human-written holdout and a machine-written holdout and report both. That one change is the paper’s most transferable finding, and it costs nothing but the labelling.