cd/entity/AdamW· home entities AdamW
grep -l @adamw /news/*.json | wc -l → 26

AdamW

mentions 26 type Organization page 1/2 feed RSS

// recent coverage 26 mentions

04:00
2026-08-31
machinebrief.com
machine-learning

Blog: Survey of Optimizers

A new survey of neural-network optimizers from arXiv (arXiv:2608.28557v1) finds that matrix-aware methods such as Muon, Shampoo, and SOAP represent a genuine advance, but no context-independent replac…

10:55
2026-08-02
promptcube3.com
large-language-models

How Much VRAM to Fine-Tune an LLM? 12 to 120 GB

Fine-tuning a 7B-parameter LLM requires 12 to 120 GB of VRAM depending on the method, according to a practical guide. Full fine-tuning in fp16 needs 80–120 GB, LoRA needs 24–32 GB, QLoRA needs 12–16 G…

19:34
2026-07-31
idlemachines.co.uk
machine-learning

Adam and AdamW: adaptive optimisation and weight decay

In a technical explainer, the author demonstrates that L2 regularization and weight decay, equivalent in plain SGD, diverge under the Adam optimizer because Adam's adaptive step rescales the L2 penalt…

04:00
2026-07-24
machinebrief.com
large-language-models

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Researchers at NVIDIA have enhanced preconditioned gradient methods like SOAP and Muon to overcome computational cost and numerical stability challenges in large-scale LLM pretraining, demonstrating t…

04:00
2026-07-24
arxiv.org
machine-learning

The Active Ingredient in Muon's Grokking

A new study from researchers at arXiv finds that the Muon optimizer's speedup in reaching the grokking threshold on modular arithmetic comes from orthogonalization via the Newton-Schulz iteration, not…

04:55
2026-07-16
thegustafson.com
machine-learning

Optimizers: Momentum, Adam, and Learning Rate Schedules

A technical blog post explains that AdamW with warmup and cosine decay is the optimizer almost everyone uses today to train large language models, tracing the evolution from vanilla SGD through moment…

04:00
2026-07-16
arxiv.org
machine-learning

Reassessing Muon for Matrix Factorization

A new study on arXiv (2607.13246v1) finds that Muon, an optimizer that reshapes gradient updates through approximate orthogonalization and has outperformed Adam and AdamW in large language model train…

07:38
2026-07-11
machinebrief.com
artificial-intelligence

The Symmetrical World of Equivariant Models: A New Era in AI

Researchers unveiled a latent world model using an equivariant encoder and predictor that achieves invariant predictions across orientations, with training loss symmetry ensuring consistent performanc…

05:13
2026-07-01
ayushtambde.com
machine-learning

Matrix Orthogonalization Improves Memory in Recurrent Models

Researchers improved associative recall in recurrent neural networks by orthogonalizing the mLSTM memory matrix during reads, inspired by the Muon optimizer. In noisy associative recall tasks, the met…

16:38
2026-06-29
lesswrong.com
large-language-models

Gradient-free Single-pass Model Beats nanoGPT on Shakespeare

A new character-level language model called EntropyBeam, using gradient-free count tables and a Dirichlet prior, achieved a validation loss of 1.596 nats on the Shakespeare character benchmark, outper…

18:02
2026-06-23
discuss.huggingface.co
machine-learning

Looking for arXiv cs.LG endorser for PsiLogic optimizer

Ali, a 16-year-old independent researcher, developed PsiLogic, a chaos-aware optimizer built on Adam, and seeks an arXiv cs.LG endorser to submit his preprint. His benchmarks on an NVIDIA H100 GPU sho…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics