Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory Researchers introduced Spatial Memory Intelligence (SMI), a framework that uses a multimodal large language model to manage spatial memory in long-video world models, according to arXiv paper 2610.02521v1. SMI coordinates four atomic operations — spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering — and experiments across multiple baselines, benchmarks, and world-model backbones showed improvements in memory sparsity, generation stability, and spatial consistency. arXiv:2610.02521v1 Announce Type: new Abstract: Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observations conditioned on user actions and historical memory. However, as memory sequences grow longer and their structures become increasingly complex, managing long-range spatial context becomes increasingly challenging, calling for a more intelligent and systematic memory-management strategy. Building on the advancing spatial reasoning capabilities of multimodal large language models MLLMs and the broader vision of unified models, we propose Spatial Memory Intelligence SMI , the first framework to systematically employ an understanding model for spatial-memory management in long-video world models. SMI introduces four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Extensive experiments across multiple baselines, benchmarks, and world-model backbones demonstrate the effectiveness and generalizability of SMI, achieving comprehensive improvements in memory sparsity, generation stability, and spatial consistency.