# VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

> Source: <https://aiflash.com/news/121135/>
> Published: 2026-09-17 03:30:19+00:00

Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kernel must be accurate and fast. Accuracy is limited by outliers: a block's quantization scale is set by its largest entrie
