{"slug": "criticality-in-dissimilar-decomposition-and-undersampling-of-random-datasets", "title": "Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies", "summary": "A new arXiv paper (2609.13201v1) models AI-generated text and images as anomalies linked to main data points and derives bounds on the minimum size of a strongly dissimilar decomposition of random datasets using redundancy graphs and iteration techniques. The authors report a phase transition in which that minimum size is determined by the main data points when anomalies are few but is taken over by the anomalies above a certain threshold, and they establish a size criticality result for strong similarity of a randomly undersampled dataset. The work is framed around the effect of AI-generated data on batch decompositions and the performance of future LLMs trained on such datasets.", "body_md": "arXiv:2609.13201v1 Announce Type: new \nAbstract: Training datasets for upcoming LLMs would include a significant amount of AI text/image data generated from current LLMs. In such a scenario, it is important to understand how this affects batch decompositions and thereby, the performance of the resultant new LLM. In this paper, we consider AI generated data as anomalies ``linked\" to main data points and study decomposition and undersampling properties of the overall random dataset. We use redundancy graphs and iteration techniques to obtain bounds for the minimum size of a strongly dissimilar (SD) decomposition and demonstrate a phase transition phenomena, wherein the minimum size is essentially determined by the \\emph{main} data points when the number of anomalies is small and is ``taken\" over by the anomalies above a certain threshold. We also establish a size criticality result for the strong similarity of a randomly undersampled dataset and illustrate our results with examples involving categorical datasets, whose overall space size is much larger than the size of the dataset.", "url": "https://wpnews.pro/news/criticality-in-dissimilar-decomposition-and-undersampling-of-random-datasets", "canonical_source": "https://arxiv.org/abs/2609.13201", "published_at": "2026-09-15 04:00:00+00:00", "updated_at": "2026-09-15 04:30:19.175532+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "ai-research"], "entities": ["arXiv", "LLMs"], "alternates": {"html": "https://wpnews.pro/news/criticality-in-dissimilar-decomposition-and-undersampling-of-random-datasets", "markdown": "https://wpnews.pro/news/criticality-in-dissimilar-decomposition-and-undersampling-of-random-datasets.md", "text": "https://wpnews.pro/news/criticality-in-dissimilar-decomposition-and-undersampling-of-random-datasets.txt", "jsonld": "https://wpnews.pro/news/criticality-in-dissimilar-decomposition-and-undersampling-of-random-datasets.jsonld"}}