19:34
2026-07-31
idlemachines.co.uk
machine-learning
Adam and AdamW: adaptive optimisation and weight decay
In a technical explainer, the author demonstrates that L2 regularization and weight decay, equivalent in plain SGD, diverge under the Adam optimizer because Adam's adaptive step rescales the L2 penaltβ¦