cd /news/machine-learning/curvature-conditioned-multiscale-mom… · home topics machine-learning article
[ARTICLE · art-116257] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining

Researchers propose a curvature-conditioned multiscale momentum method with sphere constraints that accelerates LLM pretraining, improving Muon across dense and MoE architectures from 0.12B to 2.3B parameters. The method targets flat directions to speed up final loss reduction, addressing ill-conditioned loss landscapes.

read1 min views1 publishedAug 31, 2026

arXiv:2608.28442v1 Announce Type: new Abstract: Pretraining accounts for a large fraction of the total computational cost in LLM training. However, noise-dominant gradients and the highly ill-conditioned loss landscape bring severe challenges. Although modern adaptive optimizers such as AdamW and Muon have achieved great success in large-scale pretraining, their reliance on gradient normalization offers limited mitigation of the ill-conditioned curvature. The progress along flat directions (eigen-directions of small eigenvalues), which dominates the final loss reduction, remains relatively slow. To enhance training dynamics along flat directions, we propose a curvature-conditioned multiscale momentum method with sphere constraints, delivering steady acceleration in LLM pretraining. This multiscale momentum, applied only along flat directions, pairs a slow-decay component for noise reduction with a fast-decay component for rapid curvature adaptation, harnessing their complementary strengths. Crucially, we employ a sphere constraint technique to prevent parameter inflation and excessively rapid effective learning rate decay that would otherwise arise from a naive combination. Extensive experiments show that the proposed method significantly accelerates Muon across diverse architectures (dense, MoE) and model sizes (0.12B--2.3B parameters). Theoretically, we verify the acceleration effect and provide insight into the design principles underlying the flat-direction multiscale momentum.

── more in #machine-learning 4 stories · sorted by recency
── more on @muon 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/curvature-conditione…] indexed:0 read:1min 2026-08-31 ·