Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration
On NVIDIA Blackwell B200 GPUs, tensor-core throughput exceeds special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation as a bottleneck in fused att…