A note on $L^2 \leftrightarrow L^1$ duality in shallow homogeneous nets A new technical note explains that in a shallow homogeneous neural net, minimizing the L2 norm of the parameters is equivalent to minimizing the L1 norm of a measure over unit-norm neurons, a duality the author says is directly relevant to sparse autoencoders (SAEs) and their tendency toward sparsity. The note derives the equivalence by reparameterizing a general matrix factorization F = AB^T, showing that at optimum the row norms of A and B are equal and that the inner optimization reduces to a sum of nonnegative scale variables s_i over rank-one unit-norm matrices. If you’ve been around the deep learning theory block long enough to get to know the locals, you’ve probably heard something like this at some point: In a shallow homogeneous neural net, minimizing the $L^2$-norm of the parameters is equivalent to minimizing the $L^1$-norm of a measure over unit-norm neurons. This makes a link between “aggregate $L^2$ land” and “sparse $L^1$ land.” It’s also directly relevant to SAEs, which are pretty much shallow homogeneous nets, and it’s the first place you should start when trying to understand why they tend towards sparsity. As far as I know, incredibly, this project has not yet been done. @The entire field of mechinterp. …anyways, Dhruva brought this up recently and asked me to write something explaining the idea. I’ve explained this a few times to different people, but I don’t actually know a good beginner ref for it, so I said yes. To get the idea, let’s start with a general matrix factorization problem. Let $\mathbf{A} \in \mathbb{R}^{m \times h}$ and $\mathbf{B} \in \mathbb{R}^{n \times h}$ be the two factors of a matrix $\mathbf{F} = \mathbf{AB}^\top$.