{"slug": "vc-attention-value-smoothing-and-softmax-casting-for-low-bit-attention", "title": "VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention", "summary": "Researchers introduced VC-Attention, a method combining value smoothing and softmax casting to enable low-bit attention for Diffusion Transformers used in video generation. The work targets the dominant deployment cost of attention in long spatiotemporal sequences, where accuracy is limited by outliers because a block's quantization scale is set by its largest entries.", "body_md": "Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kernel must be accurate and fast. Accuracy is limited by outliers: a block's quantization scale is set by its largest entrie", "url": "https://wpnews.pro/news/vc-attention-value-smoothing-and-softmax-casting-for-low-bit-attention", "canonical_source": "https://aiflash.com/news/121135/", "published_at": "2026-09-17 03:30:19+00:00", "updated_at": "2026-09-17 03:53:45.546831+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "ai-infrastructure"], "entities": ["VC-Attention", "Diffusion Transformers"], "alternates": {"html": "https://wpnews.pro/news/vc-attention-value-smoothing-and-softmax-casting-for-low-bit-attention", "markdown": "https://wpnews.pro/news/vc-attention-value-smoothing-and-softmax-casting-for-low-bit-attention.md", "text": "https://wpnews.pro/news/vc-attention-value-smoothing-and-softmax-casting-for-low-bit-attention.txt", "jsonld": "https://wpnews.pro/news/vc-attention-value-smoothing-and-softmax-casting-for-low-bit-attention.jsonld"}}