cd/entity/Muon· home› entities› Muon
grep -l @muon /news/*.json | wc -l → 21

Muon

mentions 21 type Organization page 1/2 feed RSS

// recent coverage 21 mentions

04:00
2026-10-08
arxiv.org
machine-learning

The Best Optimizer Depends on Batch Size

A paper submitted to arXiv on 6 Oct 2026 (arXiv:2610.08975) finds that the best optimizer for language model pretraining changes with batch size even after extensive hyperparameter tuning, and that no…

00:00
2026-10-07
mindstudio.ai
large-language-models

Aleph Alpha Kolibri 1: Germany's Open-Weight MoE Reasoning Model

Aleph Alpha released Kolibri 1, a 78-billion-parameter mixture-of-experts reasoning model for German and English, on October 3, 2026 under an Apache 2.0 license. The model activates 3.46 billion param…

01:42
2026-09-22
discuss.huggingface.co
machine-learning

The LR Death Spiral 🧨 — Muon/Dion3 Debugging Saga

An engineer preparing optimizer experiments for an upcoming PyTorch Conference talk found that the Dion3 implementation in OLMo-core broke training with a learning rate death spiral while testing Adam…

04:00
2026-08-31
machinebrief.com
machine-learning

Blog: Survey of Optimizers

A new survey of neural-network optimizers from arXiv (arXiv:2608.28557v1) finds that matrix-aware methods such as Muon, Shampoo, and SOAP represent a genuine advance, but no context-independent replac…

04:00
2026-08-24
machinebrief.com
machine-learning

RODE: A Radial-Orthogonal Decoupled Engine for Optimization

Researchers introduced RODE, a matrix-aware optimizer that decouples radial and directional updates, outperforming Muon variants across language modeling and image classification tasks. At 1.5B scale,…

00:25
2026-08-23
github.com
artificial-intelligence

NanoGPT Speedrun

The NanoGPT speedrun repository reports that a collaborative effort has trained a language model to 3.28 cross-entropy loss on the FineWeb validation set in under 75 seconds on 8 NVIDIA H100 GPUs, a d…

04:00
2026-08-17
machinebrief.com
machine-learning

Approximate Muon with low-rank adapters

Researchers propose sMuon, a method that adapts the Muon optimizer for low-rank fine-tuning by approximating the orthogonalization step via linearization and least-squares, using only matmul operation…

04:00
2026-07-24
machinebrief.com
large-language-models

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Researchers at NVIDIA have enhanced preconditioned gradient methods like SOAP and Muon to overcome computational cost and numerical stability challenges in large-scale LLM pretraining, demonstrating t…

04:00
2026-07-24
arxiv.org
machine-learning

The Active Ingredient in Muon's Grokking

A new study from researchers at arXiv finds that the Muon optimizer's speedup in reaching the grokking threshold on modular arithmetic comes from orthogonalization via the Newton-Schulz iteration, not…

04:00
2026-07-16
arxiv.org
machine-learning

Reassessing Muon for Matrix Factorization

A new study on arXiv (2607.13246v1) finds that Muon, an optimizer that reshapes gradient updates through approximate orthogonalization and has outperformed Adam and AdamW in large language model train…

07:38
2026-07-11
machinebrief.com
artificial-intelligence

The Symmetrical World of Equivariant Models: A New Era in AI

Researchers unveiled a latent world model using an equivariant encoder and predictor that achieves invariant predictions across orientations, with training loss symmetry ensuring consistent performanc…

05:13
2026-07-01
ayushtambde.com
machine-learning

Matrix Orthogonalization Improves Memory in Recurrent Models

Researchers improved associative recall in recurrent neural networks by orthogonalizing the mLSTM memory matrix during reads, inspired by the Muon optimizer. In noisy associative recall tasks, the met…

04:00
2026-06-15
arxiv.org
machine-learning

Muon$^p$: Muon with Fractional Spectral Powers

Researchers introduced Muon$^p$, a new optimizer that uses fractional spectral-power updates to interpolate between Muon and gradient descent, improving finetuning performance on billion-scale models.…

07:31
2026-05-19
latent.space
artificial-intelligence

[AINews] How to land a job at a frontier lab (on Pretraining)

Vlad Feinberg published a guide on how to land a job at a frontier AI lab focused on pretraining, emphasizing kernel-level performance work as the most direct path into the labs. The guide recommends …

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics