HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models Researchers introduced HLA-WM, a hybrid linear attention architecture for long-horizon video world models that combines softmax attention's full-history KV cache with recurrent linear attention's fixed-size state compression to cut memory use. The method targets persistent scene consistency over extended rollouts, where softmax attention's growing KV cache and linear attention's compressed fixed-size states each pose trade-offs. Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memo