I was supposed to be running optimizer experiments for my upcoming PyTorch Conference talk.
Instead, my learning rate entered a death spiral.
I was testing AdamW, Muon, Dion, and the latest Dion3 implementation in OLMo-core—and Dion3 broke things in some very interesting ways.
Turns out the debugging trail involved distributed optimizer state, LR scheduling, and a particularly nasty PyTorch tensor aliasing trap.
A fun debugging experience nonetheless.
I wrote up the whole adventure here:
More optimizer experiments coming before PyTorch Con!