arXiv:2609.09589v1 Announce Type: new
Abstract: Deep neural networks exhibit regular macroscopic behavior despite highly nonlinear dynamics in vast parameter spaces. We develop a statistical-mechanical description of learning directly in function space, treating parameter configurations as microscopic realizations and functions with their dynamical operators as macroscopic variables. For mean-squared loss, the exact error dynamics are governed by the learning operator $M=JJ^\ast$. Combining the dynamical Boltzmann weight of the conditional stochastic dynamics with the parameter-space density of states, whose local curvature defines a statistical operator $B$, and integrating over local fluctuations yields
$$ \Phi_{\mathrm{fluc}}(M;B)=\frac{\sigma_\xi^2}{2}\log\det(M^{-1}+B)+\mathrm{const}. $$
At fixed spectrum, this term is rotationally stationary when $[M,B]=0$, is minimized by pairing large eigenvalues of $M$ with small eigenvalues of $B$, and generates a local restoring contribution against rotational mismatch. For ReLU-type function spaces under mild stable statistical conditions, $B=\sigma_\xi^2L^\ast\mathcal K L$, where $L$ measures coarse-grained second-order structure. Thus the low-$B$ sector corresponds, up to bounded anisotropy of $\mathcal K$, to low structural curvature, implying a preference for faster relaxation along smooth, data-adaptive directions. These results identify function space as a natural macroscopic level for studying stable collective organization in learning.
A Function-Space Approach to the Statistical Mechanics of Learning Dynamics
A new arXiv paper (2609.09589v1) develops a statistical-mechanical description of deep neural network learning directly in function space, treating parameter configurations as microscopic realizations and functions with their dynamical operators as macroscopic variables. For mean-squared loss, the authors show the exact error dynamics are governed by the learning operator M=JJ*, and combining the dynamical Boltzmann weight with the parameter-space density of states yields the fluctuation term Φ_fluc(M;B)=(σ_ξ²/2)log det(M⁻¹+B)+const. The work identifies function space as a natural macroscopic level for studying stable collective organization in learning, with the low-B sector corresponding to low structural curvature and a preference for faster relaxation along smooth, data-adaptive directions.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.