MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models Researchers proposed MoR-MLLM, a computation-sparse multimodal large language model built on the Mixture-of-Recursions (MoR) framework that uses adaptive per-token recursion to dynamically adjust recursive depth, allocating more computation to visually or linguistically challenging tokens and skipping redundant operations for simpler ones. The work, published as arXiv:2610.08830v1, adds a three-stage MoR-Tuning strategy and an entropy-regularized loss to stabilize recursive sparsity training in multimodal settings. Experiments show MoR-MLLM greatly reduces training memory and computation complexity compared with recent advanced tiny MLLMs while retaining high performance on various vision-language tasks. arXiv:2610.08830v1 Announce Type: new Abstract: Multimodal Large Language Models MLLMs have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent efforts reduce costs by employing lightweight language backbones, existing paradigms remain computation-dense due to their static sparsity and depth allocation, which cannot adapt to the semantic complexity of each token. To this end, we propose MoR-MLLM, a computation-sparse MLLM based on the recent Mixture-of-Recursions MoR framework. MoR-MLLM introduces adaptive per-token recursion, allowing the model to dynamically adjust its recursive depth and allocate more computation to visually or linguistically challenging tokens while skipping redundant operations for simpler ones. To stabilize the training of recursive sparsity in multimodal settings, we further design a three-stage MoR-Tuning strategy and an entropy-regularized loss to encourage diverse routing distributions. Extensive experiments show that compared with recent advanced tiny MLLMs, our proposed MoR-MLLM can greatly reduce the training memory and computation complexity while retaining high performance on various vision-language tasks.