FP8 Training on AMD GPUs with TorchTitan and TorchAO: Upstreaming Performance Improvements
AMD and PyTorch upstreamed FP8 training optimizations for AMD Instinct GPUs into TorchAO and TorchTitan, delivering a 13.4% throughput gain over BF16 on Llama3-8B dense models and recovering 89% of FP…