cd /news/artificial-intelligence/mxattention-data-free-optimal-scalin… · home topics artificial-intelligence article
[ARTICLE · art-76412] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention

Researchers propose MXAttention, a data-free post-training quantization framework for MXFP4 attention that introduces Universal Optimal Scaling (UOS) with a distribution-independent optimal scaling boundary Qmax=7.25 and Pre-Normalization Quantization (PNQ) to address numerical issues in MXFP4 quantization. Experiments on Wan2.2 and HunyuanVideo show MXAttention closes at least 95% of the VBench Imaging Quality gap between OCP MXFP4 and FP16, preserving FP16-level generation quality with less than 0.01 absolute degradation on all reported VBench metrics.

read1 min views1 publishedJul 28, 2026

arXiv:2607.24377v1 Announce Type: cross Abstract: The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numerical issues: the clipping-underflow trade-off from power-of-two scaling and the row-wise normalization error introduced in the softmax loop. We propose MXAttention, a data-free post-training quantization framework for MXFP4 attention. MXAttention introduces two components: Universal Optimal Scaling (UOS), which exploits the periodic structure of power-of-two microscaling to derive a distribution-independent optimal scaling boundary Qmax=7.25 without calibration or search, and Pre-Normalization Quantization (PNQ), which quantizes unnormalized softmax exponentials before row-wise summation to preserve normalization by construction. Experiments on Wan2.2 and HunyuanVideo show that MXAttention closes at least 95% of the VBench Imaging Quality gap between OCP MXFP4 and FP16, substantially improves frame-level similarity, and preserves FP16-level generation quality with less than 0.01 absolute degradation on all reported VBench metrics. MXAttention also achieves performance competitive with strong NVFP4-based baselines with negligible overhead when fused into the attention pipeline. The implementation is publicly available in MindIE-SD.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @mxattention 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mxattention-data-fre…] indexed:0 read:1min 2026-07-28 ·