cd /news/computer-vision/pro-bench-prompt-robust-open-vocabul… · home topics computer-vision article
[ARTICLE · art-138782] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=· neutral

Pro-Bench: Prompt-Robust Open-Vocabulary Visual Grounding Across Real-World Heterogeneous Environments

Researchers introduced Pro-Bench, a prompt-conditioned benchmark for open-vocabulary visual grounding in heterogeneous real-world environments, comprising 13k+ RGB frames from subterranean, industrial, indoor, outdoor and urban robotic domains, 74.5k manual instance annotations and 515 target queries. Benchmarking 16 open-vocabulary model configurations in strict zero-shot inference, the team found prompt-robustness is strongly architecture-dependent: 10 of 16 configurations perform best with short category labels, while free-form queries yield the highest accuracy for only one, and similar aggregate mAP can conceal substantial differences in consistent target recovery across reformulations.

by read1 min views2 publishedSep 24, 2026

arXiv:2609.27076v1 Announce Type: new Abstract: Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perceptual taxonomies. However, existing benchmarks largely rely on short category labels and web-scraped imagery, leaving it unclear whether open-vocabulary models can robustly ground diverse queries and visual conditions under real deployments. We introduce \textbf{Pro-Bench}, a prompt-conditioned benchmark for open-vocabulary visual grounding in heterogeneous, real-world environments. Pro-Bench includes $13k+$ RGB frames from independent robotic domains (subterranean, industrial, indoor, outdoor, urban), with $74.5k$ manual instance annotations and $515$ target queries covering categorical, attributive, relational, affordance, state, part-whole, negative, and compositional semantics. We benchmarked $16$ open-vocabulary model configurations in strict zero-shot inference, measuring localisation accuracy across IoU thresholds, end-to-end inference latency, prompt-induced performance variation, and target recovery consistency. Our results show that prompt-robustness is strongly architecture-dependent. Most model configurations ($10/16$) perform best with short category labels, whereas free-form queries yield the highest accuracy for only one. Moreover, similar aggregate mAP can conceal substantial differences in consistent target recovery across reformulations. Pro-Bench enables systematic evaluation of these gaps and supports prompt-robust visual grounding. Pro-Bench: https://pro-bench.github.io/.

── more in #computer-vision 4 stories · sorted by recency
── more on @pro-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/pro-bench-prompt-rob…] indexed:0 read:1min 2026-09-24 ·