cd /news/machine-learning/layerwise-decoupling-for-stable-stru… · home topics machine-learning article
[ARTICLE · art-135532] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers

A new arXiv paper (arXiv:2609.21126v1) proposes a decoupled, layerwise method for structurally sparsifying the fully connected layers of pretrained neural networks, extracting shallow two-layer subnetworks, normalizing inner weights, and applying a structured group penalty to each block's outer weight matrix sequentially. The authors prove the constrained decoupled objective is equivalent at optimality to a specific joint penalty on inner and outer weights for any positively homogeneous activation, and report that the decoupled reformulation offers a wider usable range of regularization strength and a lower rate of catastrophic over-pruning than the tested joint baseline while maintaining comparable accuracy. The properties are established in controlled classification and sparse-recovery studies and examined in a high-dimensional PINN stress test and the feed-forward layers of OPT-1.3B.

by read1 min views1 publishedSep 21, 2026

arXiv:2609.21126v1 Announce Type: new Abstract: We propose a decoupled, layerwise method for structurally sparsifying the fully connected layers of pretrained neural networks. Rather than penalizing all layers jointly, our approach extracts shallow two-layer subnetworks, normalizes the inner weights, and applies a structured group penalty to the outer weight matrix of each block, processing layers sequentially to prune neurons and reduce the width of each layer. We prove that the constrained decoupled objective is equivalent at optimality to a specific joint penalty on the inner and outer weights, for any positively homogeneous activation, and thus admits a clean projected and proximal formulation. Our central finding is that this decoupled reformulation is more robust than coupled methods. In numerical experiments it provides a wider usable range of the regularization strength and a lower rate of catastrophic over-pruning than the tested joint baseline while maintaining comparable accuracy. We establish these properties in controlled classification and sparse-recovery studies, and examine their scope in a high-dimensional PINN stress test and in the feed-forward layers of OPT-1.3B.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/layerwise-decoupling…] indexed:0 read:1min 2026-09-21 ·