cd/entity/Muon· home entities Muon
grep -l @muon /news/*.json | wc -l → 13

Muon

mentions 13 type Organization feed RSS

// recent coverage 13 mentions

00:25
2026-08-23
github.com
artificial-intelligence

NanoGPT Speedrun

The NanoGPT speedrun repository reports that a collaborative effort has trained a language model to 3.28 cross-entropy loss on the FineWeb validation set in under 75 seconds on 8 NVIDIA H100 GPUs, a d…

04:00
2026-08-17
machinebrief.com
machine-learning

Approximate Muon with low-rank adapters

Researchers propose sMuon, a method that adapts the Muon optimizer for low-rank fine-tuning by approximating the orthogonalization step via linearization and least-squares, using only matmul operation…

04:00
2026-07-24
arxiv.org
machine-learning

The Active Ingredient in Muon's Grokking

A new study from researchers at arXiv finds that the Muon optimizer's speedup in reaching the grokking threshold on modular arithmetic comes from orthogonalization via the Newton-Schulz iteration, not…

04:00
2026-07-24
machinebrief.com
large-language-models

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Researchers at NVIDIA have enhanced preconditioned gradient methods like SOAP and Muon to overcome computational cost and numerical stability challenges in large-scale LLM pretraining, demonstrating t…

04:00
2026-07-16
arxiv.org
machine-learning

Reassessing Muon for Matrix Factorization

A new study on arXiv (2607.13246v1) finds that Muon, an optimizer that reshapes gradient updates through approximate orthogonalization and has outperformed Adam and AdamW in large language model train…

07:38
2026-07-11
machinebrief.com
artificial-intelligence

The Symmetrical World of Equivariant Models: A New Era in AI

Researchers unveiled a latent world model using an equivariant encoder and predictor that achieves invariant predictions across orientations, with training loss symmetry ensuring consistent performanc…

05:13
2026-07-01
ayushtambde.com
machine-learning

Matrix Orthogonalization Improves Memory in Recurrent Models

Researchers improved associative recall in recurrent neural networks by orthogonalizing the mLSTM memory matrix during reads, inspired by the Muon optimizer. In noisy associative recall tasks, the met…

04:00
2026-06-15
arxiv.org
machine-learning

Muon$^p$: Muon with Fractional Spectral Powers

Researchers introduced Muon$^p$, a new optimizer that uses fractional spectral-power updates to interpolate between Muon and gradient descent, improving finetuning performance on billion-scale models.…

07:31
2026-05-19
latent.space
artificial-intelligence

[AINews] How to land a job at a frontier lab (on Pretraining)

Vlad Feinberg published a guide on how to land a job at a frontier AI lab focused on pretraining, emphasizing kernel-level performance work as the most direct path into the labs. The guide recommends …

// co-occurs with top 8 entities
// topics top 6 topics