{"slug": "mxattention-data-free-optimal-scaling-and-pre-normalization-quantization-for", "title": "MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention", "summary": "Researchers propose MXAttention, a data-free post-training quantization framework for MXFP4 attention that introduces Universal Optimal Scaling (UOS) with a distribution-independent optimal scaling boundary Qmax=7.25 and Pre-Normalization Quantization (PNQ) to address numerical issues in MXFP4 quantization. Experiments on Wan2.2 and HunyuanVideo show MXAttention closes at least 95% of the VBench Imaging Quality gap between OCP MXFP4 and FP16, preserving FP16-level generation quality with less than 0.01 absolute degradation on all reported VBench metrics.", "body_md": "arXiv:2607.24377v1 Announce Type: cross\nAbstract: The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numerical issues: the clipping-underflow trade-off from power-of-two scaling and the row-wise normalization error introduced in the softmax loop. We propose MXAttention, a data-free post-training quantization framework for MXFP4 attention. MXAttention introduces two components: Universal Optimal Scaling (UOS), which exploits the periodic structure of power-of-two microscaling to derive a distribution-independent optimal scaling boundary Qmax=7.25 without calibration or search, and Pre-Normalization Quantization (PNQ), which quantizes unnormalized softmax exponentials before row-wise summation to preserve normalization by construction. Experiments on Wan2.2 and HunyuanVideo show that MXAttention closes at least 95% of the VBench Imaging Quality gap between OCP MXFP4 and FP16, substantially improves frame-level similarity, and preserves FP16-level generation quality with less than 0.01 absolute degradation on all reported VBench metrics. MXAttention also achieves performance competitive with strong NVFP4-based baselines with negligible overhead when fused into the attention pipeline. The implementation is publicly available in MindIE-SD.", "url": "https://wpnews.pro/news/mxattention-data-free-optimal-scaling-and-pre-normalization-quantization-for", "canonical_source": "https://www.machinebrief.com/news/mxattention-data-free-optimal-scaling-and-pre-normalization-t7gb", "published_at": "2026-07-28 04:00:00+00:00", "updated_at": "2026-07-28 04:56:41.950473+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-infrastructure"], "entities": ["MXAttention", "Universal Optimal Scaling", "Pre-Normalization Quantization", "Wan2.2", "HunyuanVideo", "VBench", "MindIE-SD", "MXFP4"], "alternates": {"html": "https://wpnews.pro/news/mxattention-data-free-optimal-scaling-and-pre-normalization-quantization-for", "markdown": "https://wpnews.pro/news/mxattention-data-free-optimal-scaling-and-pre-normalization-quantization-for.md", "text": "https://wpnews.pro/news/mxattention-data-free-optimal-scaling-and-pre-normalization-quantization-for.txt", "jsonld": "https://wpnews.pro/news/mxattention-data-free-optimal-scaling-and-pre-normalization-quantization-for.jsonld"}}