{"slug": "exploiting-target-knowledge-from-mllms-for-robust-few-shot-segmentation", "title": "Exploiting Target Knowledge from MLLMs for Robust Few-Shot Segmentation", "summary": "A new framework called MK-FSS mines target knowledge from Multimodal Large Language Models to improve few-shot segmentation, according to an arXiv paper (arXiv:2609.28949v1). Built on SAM 2, MK-FSS combines spatial knowledge, encoded as a memory representation and fused with support-guided memory via a dual-memory debate-fusion module, with semantic knowledge encoded as text and fused through a progressive cross-modal prompt generator. The authors report MK-FSS largely surpasses existing methods in extensive experiments, with code to be released.", "body_md": "arXiv:2609.28949v1 Announce Type: new \nAbstract: Few-shot segmentation (FSS) aims to segment unseen object categories with a few (e.g., one or five) labeled examples, enabling efficient adaptation to novel classes. Conventional models typically rely on appearance-based visual matching between support and query images for segmentation. While straightforward, these methods often struggle to handle significant appearance discrepancies and occlusions in the query image due to insufficient target knowledge. To mitigate this, we introduce a novel framework that mines target knowledge using the strong reasoning capacity of Multimodal Large Language Models (MLLMs) and employs it to enhance FSS. Specifically, building on SAM 2, our method, named MK-FSS, exploits two forms of complementary knowledge derived from a query image by an MLLM for FSS, including spatial knowledge, which provides a spatial prior indicating the potential target location, and semantic knowledge, which describes the target using text. The spatial knowledge is first encoded into a memory representation, and then resulting memory is integrated with the support-guided memory feature from query image through a carefully designed dual-memory debate-fusion (DMDF) module, yielding a more robust target memory feature. In parallel, the semantic knowledge is encoded into the textual feature, which is fused with multi-scale query features via a progressive cross-modal prompt generator (PCPG), producing a target-aware multimodal prompt for segmentation. Working together, the dual-memory feature and the multimodal prompt provide a comprehensive representation of the target, enabling more robust segmentation. In our extensive experiments, MK-FSS shows promising results and largely surpasses existing methods. Code will be released.", "url": "https://wpnews.pro/news/exploiting-target-knowledge-from-mllms-for-robust-few-shot-segmentation", "canonical_source": "https://arxiv.org/abs/2609.28949", "published_at": "2026-09-25 04:00:00+00:00", "updated_at": "2026-09-25 04:01:53.865376+00:00", "lang": "en", "topics": ["computer-vision", "large-language-models", "machine-learning", "ai-research"], "entities": ["MK-FSS", "SAM 2", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/exploiting-target-knowledge-from-mllms-for-robust-few-shot-segmentation", "markdown": "https://wpnews.pro/news/exploiting-target-knowledge-from-mllms-for-robust-few-shot-segmentation.md", "text": "https://wpnews.pro/news/exploiting-target-knowledge-from-mllms-for-robust-few-shot-segmentation.txt", "jsonld": "https://wpnews.pro/news/exploiting-target-knowledge-from-mllms-for-robust-few-shot-segmentation.jsonld"}}