Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models of
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness