{"slug": "layerwise-decoupling-for-stable-structured-sparsification-of-fully-connected", "title": "Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers", "summary": "A new arXiv paper (arXiv:2609.21126v1) proposes a decoupled, layerwise method for structurally sparsifying the fully connected layers of pretrained neural networks, extracting shallow two-layer subnetworks, normalizing inner weights, and applying a structured group penalty to each block's outer weight matrix sequentially. The authors prove the constrained decoupled objective is equivalent at optimality to a specific joint penalty on inner and outer weights for any positively homogeneous activation, and report that the decoupled reformulation offers a wider usable range of regularization strength and a lower rate of catastrophic over-pruning than the tested joint baseline while maintaining comparable accuracy. The properties are established in controlled classification and sparse-recovery studies and examined in a high-dimensional PINN stress test and the feed-forward layers of OPT-1.3B.", "body_md": "arXiv:2609.21126v1 Announce Type: new \nAbstract: We propose a decoupled, layerwise method for structurally sparsifying the fully connected layers of pretrained neural networks. Rather than penalizing all layers jointly, our approach extracts shallow two-layer subnetworks, normalizes the inner weights, and applies a structured group penalty to the outer weight matrix of each block, processing layers sequentially to prune neurons and reduce the width of each layer. We prove that the constrained decoupled objective is equivalent at optimality to a specific joint penalty on the inner and outer weights, for any positively homogeneous activation, and thus admits a clean projected and proximal formulation. Our central finding is that this decoupled reformulation is more robust than coupled methods. In numerical experiments it provides a wider usable range of the regularization strength and a lower rate of catastrophic over-pruning than the tested joint baseline while maintaining comparable accuracy. We establish these properties in controlled classification and sparse-recovery studies, and examine their scope in a high-dimensional PINN stress test and in the feed-forward layers of OPT-1.3B.", "url": "https://wpnews.pro/news/layerwise-decoupling-for-stable-structured-sparsification-of-fully-connected", "canonical_source": "https://www.machinebrief.com/news/layerwise-decoupling-for-stable-structured-sparsification-of-wsqp", "published_at": "2026-09-21 04:00:00+00:00", "updated_at": "2026-09-21 04:24:59.441940+00:00", "lang": "en", "topics": ["machine-learning", "neural-networks", "ai-research", "large-language-models"], "entities": ["arXiv", "OPT-1.3B"], "alternates": {"html": "https://wpnews.pro/news/layerwise-decoupling-for-stable-structured-sparsification-of-fully-connected", "markdown": "https://wpnews.pro/news/layerwise-decoupling-for-stable-structured-sparsification-of-fully-connected.md", "text": "https://wpnews.pro/news/layerwise-decoupling-for-stable-structured-sparsification-of-fully-connected.txt", "jsonld": "https://wpnews.pro/news/layerwise-decoupling-for-stable-structured-sparsification-of-fully-connected.jsonld"}}