{"slug": "reasoning-guided-part-level-visual-grounding-via-reinforcement-learning", "title": "Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning", "summary": "A new method called Object-Part Hierarchical Reflective Grounding (OP-HRG) improves part-level visual grounding in multimodal large language models by introducing a coarse-to-fine reasoning pipeline that first localizes the parent object and then the part, outperforming 7B grounding LLMs and SAM3 on PascalPart, PartImageNet, and InstructPart with a 4B model.", "body_md": "arXiv:2607.15374v1 Announce Type: new\nAbstract: Multimodal large language models (MLLMs) ground whole objects well from free-form language queries, but they struggle when the query names a part rather than the object. We trace this to a missing object-part hierarchy, since parts are localized in the same single step used for objects. We propose Object-Part Hierarchical Reflective Grounding (OP-HRG), a coarse-to-fine reasoning-guided grounding strategy that first localizes the parent object and then the part within it. A self-check then reflects on the result, with an extension to re-encode the predicted crop to inspect the region it is correcting. We introduce a part-aware GRPO framework to train our pipeline with stage-wise rewards. A 4B model trained this way outperforms 7B grounding LLMs and SAM3 across PascalPart, PartImageNet, and InstructPart, and transfers to reasoning segmentation.", "url": "https://wpnews.pro/news/reasoning-guided-part-level-visual-grounding-via-reinforcement-learning", "canonical_source": "https://arxiv.org/abs/2607.15374", "published_at": "2026-07-20 04:00:00+00:00", "updated_at": "2026-07-20 13:56:11.148166+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "computer-vision", "ai-research"], "entities": ["OP-HRG", "GRPO", "PascalPart", "PartImageNet", "InstructPart", "SAM3"], "alternates": {"html": "https://wpnews.pro/news/reasoning-guided-part-level-visual-grounding-via-reinforcement-learning", "markdown": "https://wpnews.pro/news/reasoning-guided-part-level-visual-grounding-via-reinforcement-learning.md", "text": "https://wpnews.pro/news/reasoning-guided-part-level-visual-grounding-via-reinforcement-learning.txt", "jsonld": "https://wpnews.pro/news/reasoning-guided-part-level-visual-grounding-via-reinforcement-learning.jsonld"}}