arXiv:2609.21926v1 Announce Type: new Abstract: Many machine learning systems try to explain complex data - like images or financial time series - in terms of hidden, independent factors that generated them. Recovering the true underlying factors, rather than some scrambled version of them, is the central challenge of nonlinear Independent Component Analysis (nICA). We prove identifiability (exact recovery) up to trivial ambiguities for real analytic generating functions when source probability density functions have a finite number of discontinuities in the first derivative. The Laplace distribution is the most prominent example satisfying this assumption. Our proof relies on the contrast between kinks in the source distribution and the smoothness of real analytic functions. Real analytic functions comprise a broad class of generating mechanisms, and can be approximated with Normalizing Flows or Variational Autoencoders with standard activation functions (e.g., tanh, softplus, GELU), so our result applies with minimal changes to existing training pipelines. We perform experiments on real and synthetic data with both Normalizing Flows and Variational Auto-Encoders demonstrating their identifiability properties. In experiments on CelebA data we recover several interpretable latent factors controlling unique attributes across the dataset.
Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources
A new arXiv paper (2609.21926v1) proves identifiability, meaning exact recovery up to trivial ambiguities, for real analytic generating functions in nonlinear Independent Component Analysis when source probability density functions have a finite number of discontinuities in the first derivative, a class that includes the Laplace distribution. The proof exploits the contrast between kinks in the source distribution and the smoothness of real analytic functions, and the authors note the result applies with minimal changes to existing training pipelines using Normalizing Flows or Variational Autoencoders with standard activations such as tanh, softplus, and GELU. Experiments on real and synthetic data with both Normalizing Flows and Variational Autoencoders demonstrate the identifiability properties, and on CelebA data the method recovers several interpretable latent factors controlling unique attributes across the dataset.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.