{"slug": "pointrl-learning-point-level-vision-language-grounding-from-verifiable-evidence", "title": "PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence", "summary": "PointRL, a verifiable reinforcement learning framework introduced in an arXiv paper (2608.25299v1), improves point-level vision-language grounding by converting bounding boxes, masks, and instance labels into pointing instructions with hidden verifier evidence. On PointArena, PointRL raises Qwen3.5-4B's overall accuracy from 56.11% to 65.58%, with same-backbone gains on RoboSpatial, BLINK, and Ref-Adv benchmarks.", "body_md": "arXiv:2608.25299v1 Announce Type: new\nAbstract: Vision-language models (VLMs) increasingly rely on point coordinates as a compact and executable interface for visual grounding in GUI interaction, robotic manipulation, and interactive visual systems. However, learning reliable pointing behavior remains difficult because the supervision space is inherently non-unique: many coordinates may be valid within the same target region, while multi-instance instructions require target coverage, count consistency, and duplicate suppression. This work presents PointRL, a verifiable reinforcement learning framework that learns point-level grounding from existing heterogeneous annotation evidence. PointRL converts bounding boxes, masks, and instance labels into pointing instructions, while retaining their target supports, instance membership, and set constraints as hidden verifier evidence, i.e., annotations kept outside the prompt and used by a deterministic checker to score predictions. The proposed reward evaluates parseability, point validity, instance coverage, cardinality consistency, and redundant or missing predictions. On PointArena, PointRL improves the overall accuracy of Qwen3.5-4B from 56.11% to 65.58%. Further evaluations on RoboSpatial, BLINK, and Ref-Adv show same-backbone gains on the evaluated external benchmarks, suggesting that verifiable point-level feedback may benefit spatial grounding in these settings.", "url": "https://wpnews.pro/news/pointrl-learning-point-level-vision-language-grounding-from-verifiable-evidence", "canonical_source": "https://arxiv.org/abs/2608.25299", "published_at": "2026-08-27 04:00:00+00:00", "updated_at": "2026-08-27 04:21:47.629720+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "ai-research"], "entities": ["PointRL", "Qwen3.5-4B", "PointArena", "RoboSpatial", "BLINK", "Ref-Adv", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/pointrl-learning-point-level-vision-language-grounding-from-verifiable-evidence", "markdown": "https://wpnews.pro/news/pointrl-learning-point-level-vision-language-grounding-from-verifiable-evidence.md", "text": "https://wpnews.pro/news/pointrl-learning-point-level-vision-language-grounding-from-verifiable-evidence.txt", "jsonld": "https://wpnews.pro/news/pointrl-learning-point-level-vision-language-grounding-from-verifiable-evidence.jsonld"}}