Limitations of Synthetic Data Generation in Specialized Data-Scarce Domains A new arXiv study (2608.13729v1) finds that synthetic image generation from diffusion-based models fails to consistently improve downstream classifier performance over non-generative augmentation in specialized, data-scarce domains. Across five trauma classification tasks with subject-wise splits, no generative approach beat a strong baseline, with failure modes including memorization, distributional drift, and simplified canonical instances. arXiv:2608.13729v1 Announce Type: new Abstract: Advances in diffusion-based generative models have motivated the use of synthetic image generation to alleviate data scarcity in vision tasks. While this strategy has shown promise in natural image benchmarks such as ImageNet, its effectiveness in sparse, high-variance real-world domains remains unclear. In this work, we focus on domains where images differ substantially from common image datasets and additional data are expensive to obtain. Against non-generative data augmentation baselines, we evaluate the downstream classifier performance improvements yielded by two schools of generative sparse data extension: distribution modeling and sample perturbation. Across five trauma classification tasks using subject-wise train--validation splits, no generative approach consistently outperforms a strong non-generative baseline. Feature-space analysis reveals recurring failure modes: memorization or collapse, distributional drift, and generation of visually plausible but simplified canonical instances that are easier to classify than real data.