04:25
2026-10-09
ianbarber.blog
machine-learning
FA3 had some bugs
Researchers found two bugs in FlashAttention 3 (FA3) that caused gradient norm to grow a thousandfold and loss to end 0.2 nats above FP32 attention when pretraining a 450M-parameter transformer on 50B…