The Best Optimizer Depends on Batch Size
A paper submitted to arXiv on 6 Oct 2026 (arXiv:2610.08975) finds that the best optimizer for language model pretraining changes with batch size even after extensive hyperparameter tuning, and that no…
A paper submitted to arXiv on 6 Oct 2026 (arXiv:2610.08975) finds that the best optimizer for language model pretraining changes with batch size even after extensive hyperparameter tuning, and that no…
Aleph Alpha released Kolibri 1, a 78-billion-parameter mixture-of-experts reasoning model for German and English, on October 3, 2026 under an Apache 2.0 license. The model activates 3.46 billion param…
Keller Jordan published Muon, an orthogonalized optimizer short for MomentUm Orthogonalized by Newton–Schulz, as a blog post in December 2024 after it cut the CIFAR-10 training record from 3.3 to 2.6 …
Developer Shrijith Venkatramana explains Muon, an optimizer that treats neural network weight matrices as matrices rather than collections of scalar coordinates, orthogonalizing momentum updates by pr…
An engineer preparing optimizer experiments for an upcoming PyTorch Conference talk found that the Dion3 implementation in OLMo-core broke training with a learning rate death spiral while testing Adam…
A new survey of neural-network optimizers from arXiv (arXiv:2608.28557v1) finds that matrix-aware methods such as Muon, Shampoo, and SOAP represent a genuine advance, but no context-independent replac…
Researchers propose a curvature-conditioned multiscale momentum method with sphere constraints that accelerates LLM pretraining, improving Muon across dense and MoE architectures from 0.12B to 2.3B pa…
Researchers introduced RODE, a matrix-aware optimizer that decouples radial and directional updates, outperforming Muon variants across language modeling and image classification tasks. At 1.5B scale,…
The NanoGPT speedrun repository reports that a collaborative effort has trained a language model to 3.28 cross-entropy loss on the FineWeb validation set in under 75 seconds on 8 NVIDIA H100 GPUs, a d…
Researchers propose sMuon, a method that adapts the Muon optimizer for low-rank fine-tuning by approximating the orthogonalization step via linearization and least-squares, using only matmul operation…
Researchers at NVIDIA have enhanced preconditioned gradient methods like SOAP and Muon to overcome computational cost and numerical stability challenges in large-scale LLM pretraining, demonstrating t…
A new study from researchers at arXiv finds that the Muon optimizer's speedup in reaching the grokking threshold on modular arithmetic comes from orthogonalization via the Newton-Schulz iteration, not…
A new study on arXiv (2607.13246v1) finds that Muon, an optimizer that reshapes gradient updates through approximate orthogonalization and has outperformed Adam and AdamW in large language model train…
Researchers unveiled a latent world model using an equivariant encoder and predictor that achieves invariant predictions across orientations, with training loss symmetry ensuring consistent performanc…
Efficient Long-hOrizon (ELO) learning, a new learned optimizer, surpasses traditional models like AdamW and Muon across diverse tasks including language modeling and image classification, requiring le…
Researchers improved associative recall in recurrent neural networks by orthogonalizing the mLSTM memory matrix during reads, inspired by the Muon optimizer. In noisy associative recall tasks, the met…
Researchers introduce Depth-wise Gradient Augmentation, a new optimization paradigm that transforms layer-wise updates to exploit structured relationships across deep neural network layers. Their inst…
Researchers introduced Muon$^p$, a new optimizer that uses fractional spectral-power updates to interpolate between Muon and gradient descent, improving finetuning performance on billion-scale models.…
Vlad Feinberg published a guide on how to land a job at a frontier AI lab focused on pretraining, emphasizing kernel-level performance work as the most direct path into the labs. The guide recommends …
DeepSeek-V4's million-token context capability stems from a hybrid attention architecture that compresses context before KV storage, reducing cache pressure. Together's early bring-up on NVIDIA HGX B2…