{"slug": "matrix-adagrad-row-wise-and-column-wise-adaptive-subgradient-methods", "title": "Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods", "summary": "Researchers developed Row-wise Matrix AdaGrad (Row-AdaGrad) and Column-wise Matrix AdaGrad (Column-AdaGrad), matrix-aware adaptive optimization methods derived from an Online Mirror Descent framework with adaptive proximal functions, according to an arXiv paper (2609.21815v1). The paper establishes regret guarantees showing these matrix-aware bounds can be strictly tighter than those of entry-wise AdaGrad under structured gradients. Experiments on matrix factorization and deep neural-network training showed improved optimization stability and trainability at larger learning rates and greater network depths.", "body_md": "arXiv:2609.21815v1 Announce Type: new \nAbstract: Adaptive optimization methods such as AdaGrad and Adam are widely used in modern neural-network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware optimizers demonstrate the benefits of structured optimization, yet a general theoretical framework for deriving matrix-aware adaptivity comparable to that of AdaGrad remains lacking. In this work, we develop a general Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters, providing a principled approach to deriving matrix-aware adaptive optimization through online regret minimization. By introducing row-wise and column-wise matrix proximal functions and analyzing the resulting regret trade-off, we derive Row-wise Matrix AdaGrad (Row-AdaGrad) and Column-wise Matrix AdaGrad (Column-AdaGrad), with adaptive scaling determined by the accumulated row-wise or column-wise gradient norms. We establish regret guarantees and show that these matrix-aware bounds can be strictly tighter than those of entry-wise AdaGrad under structured gradients. Experiments on matrix factorization and deep neural-network training further demonstrate the benefits of aligning adaptive scaling with matrix structure, including improved optimization stability and trainability at larger learning rates and greater network depths.", "url": "https://wpnews.pro/news/matrix-adagrad-row-wise-and-column-wise-adaptive-subgradient-methods", "canonical_source": "https://www.machinebrief.com/news/matrix-adagrad-row-wise-and-column-wise-adaptive-subgradient-wyl7", "published_at": "2026-09-21 04:00:00+00:00", "updated_at": "2026-09-21 05:54:19.441154+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "artificial-intelligence", "neural-networks"], "entities": ["arXiv", "AdaGrad", "Adam", "Row-wise Matrix AdaGrad", "Column-wise Matrix AdaGrad", "Online Mirror Descent"], "alternates": {"html": "https://wpnews.pro/news/matrix-adagrad-row-wise-and-column-wise-adaptive-subgradient-methods", "markdown": "https://wpnews.pro/news/matrix-adagrad-row-wise-and-column-wise-adaptive-subgradient-methods.md", "text": "https://wpnews.pro/news/matrix-adagrad-row-wise-and-column-wise-adaptive-subgradient-methods.txt", "jsonld": "https://wpnews.pro/news/matrix-adagrad-row-wise-and-column-wise-adaptive-subgradient-methods.jsonld"}}