{"slug": "where-grounding-accuracy-lives-on-the-iou-curve-label-free-inference-time", "title": "Where Grounding Accuracy Lives on the IoU Curve: Label-Free Inference-Time Boundary Refinement", "summary": "A new arXiv paper (2608.19553v1) introduces label-free precision refinement (LFPR), a method that improves grounding bounding-box accuracy by routing predicted-small regions to a higher-resolution pass without accessing target annotations at inference. On 31,921 Ref-L4 expressions, LFPR raises mAcc0.5:0.95 from 72.947% to 76.013%, and on 30,969 RefCOCO/RefCOCO+/RefCOCOg expressions it improves every dataset at Acc@0.5, mAcc, and mean IoU, while a prospective Flickr30K Entities evaluation improves every endpoint (mAcc +0.973, Acc@0.9 +1.022). The authors show that referent selection and boundary precision are partially separable, with different components moving opposing regions of the IoU curve.", "body_md": "arXiv:2608.19553v1 Announce Type: new\nAbstract: Vision--language models can identify the correct referent while returning an imprecise bounding box. We study whether a frozen direct-answer model can use its own prediction to allocate one additional localized observation without accessing target annotations at inference. Label-free precision refinement (LFPR) routes predicted-small regions to a higher-resolution pass, re-grounds the expression inside a context crop, admits a candidate only under fixed geometric guards, and returns a fixed coordinate-wise midpoint. We report results across three evidence tiers. On 31,921 retrospective Ref-L4 expressions, LFPR raises mAcc$_{0.5:0.95}$ from 72.947\\% to 76.013\\% (Acc@0.5 88.531\\%$\\to$89.725\\%, Acc@0.9 55.788\\%$\\to$61.142\\%). A frozen transfer to 30,969 RefCOCO/RefCOCO+/RefCOCOg expressions improves every dataset at Acc@0.5, mAcc, and mean IoU (pooled mAcc $+0.645$, Acc@0.5 $+0.817$), while Acc@0.9 is unchanged overall: routing alone gains $+1.162$ points there, but crop, guards, and fusion give back $-1.192$, offsetting rather than showing no strict-IoU effect. A prospective, image-disjoint Flickr30K Entities evaluation improves every endpoint (mAcc $+0.973$, Acc@0.9 $+1.022$), more strongly under a single-box variant (mAcc $+2.575$, Acc@0.9 $+3.689$). The same operator applied to two released grounding specialists improves every endpoint (Acc@0.9 $+1.569$/$+6.716$ for EGM-4B/8B) at roughly twice the latency, composing with specialist training rather than replacing it. A genuine unguarded control (guard removed from the same candidates) underperforms the incumbent on every metric, showing the guard is load-bearing. Together, these results show that referent selection and boundary precision are partially separable, with different components moving opposing regions of the IoU curve -- behavior a single threshold cannot reveal.", "url": "https://wpnews.pro/news/where-grounding-accuracy-lives-on-the-iou-curve-label-free-inference-time", "canonical_source": "https://arxiv.org/abs/2608.19553", "published_at": "2026-08-21 04:00:00+00:00", "updated_at": "2026-08-21 04:16:49.901921+00:00", "lang": "en", "topics": ["computer-vision", "machine-learning", "artificial-intelligence"], "entities": ["arXiv", "Ref-L4", "RefCOCO", "RefCOCO+", "RefCOCOg", "Flickr30K Entities", "EGM-4B", "EGM-8B"], "alternates": {"html": "https://wpnews.pro/news/where-grounding-accuracy-lives-on-the-iou-curve-label-free-inference-time", "markdown": "https://wpnews.pro/news/where-grounding-accuracy-lives-on-the-iou-curve-label-free-inference-time.md", "text": "https://wpnews.pro/news/where-grounding-accuracy-lives-on-the-iou-curve-label-free-inference-time.txt", "jsonld": "https://wpnews.pro/news/where-grounding-accuracy-lives-on-the-iou-curve-label-free-inference-time.jsonld"}}