cd /news/machine-learning/explaining-the-saliency-map-sparsity… · home › topics › machine-learning › article
[ARTICLE · art-148036] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Explaining the Saliency Map Sparsity of Adversarially-Trained Neural Networks

A new arXiv paper (2610.10666v1) offers a theoretical explanation for why adversarially-trained neural networks produce sparse gradient saliency maps, proving that for two-layer ReLU networks, minimizers converge to a Bayes classifier with minimal gradient and Barron norm as data points, neurons, and regularization parameters scale appropriately. The authors attribute the sparsity to the anisotropic gradient norm induced by ℓ∞-attack adversarial training, which favors axis-aligned or sparse gradients, and support the theory with experiments comparing gradient ℓ1-norm and thresholded sparsity between naturally and adversarially trained models.

by read1 min views3 publishedOct 9, 2026

arXiv:2610.10666v1 Announce Type: new Abstract: Understanding why deep neural networks make a given prediction is of great importance for their safe deployment. In computer vision, saliency maps, which highlight the image region most influential for a prediction, remain a widely-used form of explanation. An empirical observation is the apparent sparsity of gradient saliency maps of adversarially-trained neural networks. In this paper, we propose a theoretical explanation of this phenomenon for two-layer ReLU networks. We build on the established equivalence of adversarial training to the minimization of the empirical risk with weight-decay penalization and an added adversarial total variation term -- valid for certain loss functions. As the number of data points and neurons grows and the regularization parameters are sent to zero at appropriate rates, we prove that minimizers converge to a Bayes classifier with minimal gradient and Barron norm. Sparsity appears since for adversarial training with $\ell_\infty$-attacks the gradient norm is anisotropic and favors axis-aligned / sparse gradients. We illustrate our theoretical findings experimentally by evaluating the gradient $\ell_1$-norm and thresholded sparsity of naturally versus adversarially trained models.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/explaining-the-salie…] indexed:0 read:1min 2026-10-09 · —