{"slug": "fast-polynomial-transcendentals-for-llms", "title": "Fast Polynomial Transcendentals for LLMs", "summary": "A new arXiv paper (2610.00049v1) reports that replacing native sigmoid, tanh, and SiLU activations with degree-3 or degree-4 bfloat16 polynomial programs improved complete training-step throughput on NVIDIA GB200 systems by 2.7% for dense SiLU, 2.9% for tanh-softcapped attention, and 8.0% for routed-expert SwiGLU, while sigmoid attention improved complete-attention forward by 7.4% and the complete GPU step by 0.3%. The isolated FP16 sweep showed speedups of 1.19–2.19x for L2-resident working sets and 1.00–1.70x for HBM-resident sets. At horizons near 100 billion tokens, final smoothed training-loss differences between polynomial and native paths ranged from -0.107 to +0.079 across the four tasks.", "body_md": "arXiv:2610.00049v1 Announce Type: new \nAbstract: Graphics processing unit (GPU) generations scale matrix, special-function, and memory pipelines at different rates, so kernel bottlenecks move as hardware evolves. FlashAttention-4 exposed this imbalance inside attention on NVIDIA Blackwell. We test whether short polynomial programs can accelerate other special-function-unit (SFU) operations in large language models (LLMs). We first compare native PyTorch evaluation with packed fused multiply--add (FMA) programs in an isolated IEEE binary16 (FP16) sweep spanning L2-resident and high-bandwidth-memory (HBM)-resident working sets. We then replace native sigmoid, tanh, and sigmoid linear unit (SiLU) with degree-3 or degree-4 bfloat16 (BF16) programs in four GB200 integration tasks: dense SiLU, tanh-softcapped attention, sigmoid attention, and routed-expert Swish-gated linear unit (SwiGLU). The programs combine analytical symmetry, target-format rounding, and packed arithmetic inside consuming kernels. The isolated paths improve by 1.19--2.19x in L2 and 1.00--1.70x in HBM. The dense-SiLU, tanh-softcapped-attention, and routed-expert substitutions improve complete training-step throughput by 2.7\\%, 2.9\\%, and 8.0\\%, respectively. The sigmoid-attention substitution improves complete-attention forward by 7.4\\% and the complete GPU step by 0.3\\%. Same-checkpoint open-weight ablations and one paired pre-training comparison per task extend the evaluation to model behavior. At common horizons near 100 billion tokens, the final smoothed training-loss differences (polynomial minus native) range from $-0.107$ to $+0.079$ across the four tasks.", "url": "https://wpnews.pro/news/fast-polynomial-transcendentals-for-llms", "canonical_source": "https://arxiv.org/abs/2610.00049", "published_at": "2026-10-03 04:00:00+00:00", "updated_at": "2026-10-03 04:07:30.089885+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-infrastructure", "machine-learning"], "entities": ["NVIDIA", "Blackwell", "GB200", "PyTorch", "FlashAttention-4", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/fast-polynomial-transcendentals-for-llms", "markdown": "https://wpnews.pro/news/fast-polynomial-transcendentals-for-llms.md", "text": "https://wpnews.pro/news/fast-polynomial-transcendentals-for-llms.txt", "jsonld": "https://wpnews.pro/news/fast-polynomial-transcendentals-for-llms.jsonld"}}