The LR Death Spiral 🧨 — Muon/Dion3 Debugging Saga An engineer preparing optimizer experiments for an upcoming PyTorch Conference talk found that the Dion3 implementation in OLMo-core broke training with a learning rate death spiral while testing AdamW, Muon, and Dion. The debugging trail traced the failure to distributed optimizer state, LR scheduling, and a PyTorch tensor aliasing trap, which the engineer documented in a write-up. More optimizer experiments are planned before PyTorch Con. I was supposed to be running optimizer experiments for my upcoming PyTorch Conference talk. Instead, my learning rate entered a death spiral. I was testing AdamW, Muon, Dion, and the latest Dion3 implementation in OLMo-core—and Dion3 broke things in some very interesting ways. Turns out the debugging trail involved distributed optimizer state, LR scheduling, and a particularly nasty PyTorch tensor aliasing trap. A fun debugging experience nonetheless. I wrote up the whole adventure here: More optimizer experiments coming before PyTorch Con