# Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration

> Source: <https://aiflash.com/news/128945/>
> Published: 2026-09-30 01:00:31+00:00

On NVIDIA Blackwell B200, tensor-core throughput outpaces special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation in fused attention kernels. A pretrained Transformer, however, may not need it evaluated accurately at every element. We characteriz
