Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration On NVIDIA Blackwell B200 GPUs, tensor-core throughput exceeds special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation as a bottleneck in fused attention kernels, according to research characterizing softmax approximation in pretrained LLMs. The work examines model sensitivity to approximating softmax and kernel acceleration, noting that a pretrained Transformer may not need the exponential evaluated accurately at every element. On NVIDIA Blackwell B200, tensor-core throughput outpaces special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation in fused attention kernels. A pretrained Transformer, however, may not need it evaluated accurately at every element. We characteriz