arXiv:2608.23938v1 Announce Type: cross Abstract: Modern score-based generative models have achieved remarkable empirical success in high-dimensional tasks such as image, audio, and video synthesis. These models reduce distribution learning to a sequence of regression problems that, if solved exactly on finite data, would ultimately reproduce the training samples. Their ability to generalize must therefore arise from the implicit or explicit regularization during training. In this work, we develop a generative counterpart to the theory of benign overfitting and algorithmic regularization for overparameterized neural networks in the supervised lazy-training regime. We study denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel. In the proportional high-dimensional regime $n\asymp d$, we derive exact risk trajectories under gradient flow training. These trajectories exhibit three phases governed by qualitatively distinct estimators: a spectral estimator that generalizes, a pure-noise score with localized peaks that interpolate the training objective, and an empirical Bayes estimator that memorizes the data. We then analyze how these estimators combine along the reverse-time SDE and characterize the distribution of the resulting samples. The analysis reveals familiar mechanisms from supervised learning, including kernel linearization and self-induced regularization from the nonlinear part of the kernel, but also reveals a distinct phenomenology specific to generative modeling.
Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime
A new arXiv preprint (2608.23938v1) develops a theory of generalization, memorization, and overfitting for diffusion models trained in the high-dimensional lazy regime, showing that denoising score matching in a vector-valued reproducing kernel Hilbert space with an inner-product kernel exhibits three distinct phases—a spectral estimator that generalizes, a pure-noise score that interpolates, and an empirical Bayes estimator that memorizes—with exact risk trajectories derived under gradient flow in the proportional regime n≍d. The analysis reveals mechanisms from supervised learning, including kernel linearization and self-induced regularization, alongside a distinct phenomenology specific to generative modeling.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.