14:26
2026-10-08
arxiv.org
machine-learning
Decentralized SGD under Heavy-Tailed Noise: Optimal Convergence Rates and the Role of Gradient Clipping
A paper submitted to arXiv on 7 October 2026 by Aleksandar Armacki shows that clipped decentralized SGD (DSGD) achieves order-optimal convergence rates under heavy-tailed noise for smooth non-convex cā¦