VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Researchers introduced VC-Attention, a method combining value smoothing and softmax casting to enable low-bit attention for Diffusion Transformers used in video generation. The work targets the dominant deployment cost of attention in long spatiotemporal sequences, where accuracy is limited by outliers because a block's quantization scale is set by its largest entries. Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kernel must be accurate and fast. Accuracy is limited by outliers: a block's quantization scale is set by its largest entrie