{"slug": "movisa-multi-token-reasoning-for-video-object-segmentation", "title": "MoVISA: Multi-Token Reasoning for Video Object Segmentation", "summary": "Researchers developed MoVISA, a multi-token reasoning method for video object segmentation that replaces the single SEG token used in Multimodal Large Language Model pipelines with multiple tokens such as SEG0 and SEG1 to track objects across frames. On the MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS benchmarks, MoVISA reports a 13.2 percent J and F improvement on MeViS and an 8.4 percent J and F improvement on ReVOS, with code and models to be released.", "body_md": "arXiv:2609.28956v1 Announce Type: new \nAbstract: Recent advances in video object segmentation with Multimodal Large Language Model (MLLM) reasoning have demonstrated the effectiveness of using a single textual token, such as SEG, to predict segmentation masks across images and videos. However, we observe that this single-token strategy lacks the granularity required to precisely localize multiple objects across time in video segmentation tasks. To address this limitation, we develop Multi-Token Reasoning for Video Object Segmentation, or MoVISA. MoVISA uses multiple segmentation tokens, such as SEG0 and SEG1, to represent an object across different frames. This design enables more fine-grained alignment between language prompts and spatio-temporal mask predictions, improving both performance and interpretability. On the challenging MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS benchmarks, our model achieves a 13.2 percent J and F improvement on MeViS and an 8.4 percent J and F improvement on ReVOS. Code and models will be released.", "url": "https://wpnews.pro/news/movisa-multi-token-reasoning-for-video-object-segmentation", "canonical_source": "https://arxiv.org/abs/2609.28956", "published_at": "2026-09-25 04:00:00+00:00", "updated_at": "2026-09-25 04:01:56.597831+00:00", "lang": "en", "topics": ["computer-vision", "artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["MoVISA", "MeViS", "DAVIS17", "ReVOS", "Ref-Youtube-VOS"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/movisa-multi-token-reasoning-for-video-object-segmentation", "markdown": "https://wpnews.pro/news/movisa-multi-token-reasoning-for-video-object-segmentation.md", "text": "https://wpnews.pro/news/movisa-multi-token-reasoning-for-video-object-segmentation.txt", "jsonld": "https://wpnews.pro/news/movisa-multi-token-reasoning-for-video-object-segmentation.jsonld"}}