cd /news/machine-learning/sphere-retraction-normalizations · home topics machine-learning article
[ARTICLE · art-87095] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Sphere Retraction Normalizations

Researchers introduced p-SpheretNorm, a one-parameter family of angular retractions for spherical residual streams in deep neural networks, showing that the exponential map used in Geodesic Normalization (GeoNorm) is just one extreme of a spectrum. On nanoGPT, all three methods (Proj-SpheretNorm, Cay-SpheretNorm, and p-SpheretNorm) outperform existing lightweight deep connection schemes, with the best validation loss achieved at finite p, indicating the exponential map is not the preferred retraction.

read1 min views1 publishedAug 5, 2026
arXiv:2608.02668v1 Announce Type: new
Abstract: Residual connections are the de facto mechanism for training deep neural networks stably. Geodesic Normalization (GeoNorm) recasts them on a Riemannian manifold, orthogonalizing each layer output against the current hidden state and applying the resulting update through the Riemannian exponential map. Every hidden state thus keeps a constant $\ell_{2}$-norm, confining the residual stream to a hypersphere. The exponential map, however, is only one member of a broad family of retraction maps. We show that on the hypersphere this entire family collapses to a single scalar design choice. What distinguishes one retraction from another is only how the magnitude of an update is converted into a rotation angle within the plane spanned by the hidden state and the update. This view places Euclidean residual connections and GeoNorm in one framework. Instantiating it with the metric projection retraction and the Cayley retraction yields Proj-SpheretNorm and Cay-SpheretNorm, which are exactly norm-preserving yet require only algebraic operations. Both prove to be members of a one-parameter family of angular retractions, $p$-SpheretNorm, whose rotation angle saturates rather than growing without bound. The two methods above are recovered exactly at $p = 1$ and $p = 2$, while the identity map and GeoNorm arise only as limits at either end. On nanoGPT, all three methods outperform existing lightweight deep connection schemes, and the best validation loss is attained at finite $p$, indicating that the exponential map is not the preferred retraction for spherical residual streams but merely one end of a spectrum.
── more in #machine-learning 4 stories · sorted by recency
── more on @geonorm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/sphere-retraction-no…] indexed:0 read:1min 2026-08-05 ·