{"slug": "multi-agent-target-existence-verification-and-learned-mask-geometry-refinement", "title": "Multi-Agent Target-Existence Verification and Learned Mask Geometry Refinement: Winning Report of the MeViS-Text Track at the 8th LSVOS Challenge 2026", "summary": "A team led by researchers from multiple institutions won the MeViS-Text track at the 8th Large-scale Video Object Segmentation (LSVOS) Challenge 2026 with their SSUPER pipeline, achieving a Final score of 0.9081339614 on the official leaderboard. The system uses three heterogeneous multimodal large language models for multi-agent target-existence verification and a training-data-only StyleRefiner for mask geometry refinement, recovering most residual no-target errors without new segmentation calls.", "body_md": "arXiv:2608.11458v1 Announce Type: new\nAbstract: We present the first-place solution to the MeViS-Text track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge 2026: referring video object segmentation guided by written motion expressions, including deceptive no-target expressions that match no object in the video and must yield empty masks in every frame. Our pipeline, SSUPER, resolves each expression into a visual concept, generates full-video candidate masklets with SAM~3.1, and selects target IDs. At every reasoning stage, three heterogeneous multimodal large language models independently execute the same stage-specific prompt before a single synthesis pass commits one schema-validated verdict. Although this system rejects every no-target expression in validation, the leaderboard reveals that a substantial share of test no-target cases still slips through. The reason is that hard negatives name a plausible object and fail only under the complete temporal predicate, so when selection and existence are decided together, a category-plausible masklet anchors the verdict. Hence, we decouple existence verification into an independent multi-agent audit of the full predicate (category, count, action, trajectory, event order, and semantic role) that distinguishes absence from temporary invisibility, discounts apparent motion caused by camera movement, and requires contradicting evidence rather than mere uncertainty for a no-target verdict. Without any new segmentation call, this audit recovers most of the residual no-target errors. A training-data-only StyleRefiner then aligns mask geometry with the annotation style of MeViSv2 while preserving every presence decision by construction, showing that once the semantics are fixed, part of the remaining error is stylistic rather than semantic. The complete system reaches a Final score of 0.9081339614 on the official challenge leaderboard.", "url": "https://wpnews.pro/news/multi-agent-target-existence-verification-and-learned-mask-geometry-refinement", "canonical_source": "https://arxiv.org/abs/2608.11458", "published_at": "2026-08-13 04:00:00+00:00", "updated_at": "2026-08-13 04:13:03.990285+00:00", "lang": "en", "topics": ["computer-vision", "large-language-models", "artificial-intelligence"], "entities": ["SSUPER", "SAM 3.1", "MeViS-Text", "LSVOS Challenge 2026", "MeViSv2", "StyleRefiner"], "alternates": {"html": "https://wpnews.pro/news/multi-agent-target-existence-verification-and-learned-mask-geometry-refinement", "markdown": "https://wpnews.pro/news/multi-agent-target-existence-verification-and-learned-mask-geometry-refinement.md", "text": "https://wpnews.pro/news/multi-agent-target-existence-verification-and-learned-mask-geometry-refinement.txt", "jsonld": "https://wpnews.pro/news/multi-agent-target-existence-verification-and-learned-mask-geometry-refinement.jsonld"}}