cd /news/computer-vision/vantage-bench-evaluating-the-infrast… · home topics computer-vision article
[ARTICLE · art-125437] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=· neutral

VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models

Researchers introduced VANTAGE-Bench, a benchmark measuring an "Infrastructure AI Gap" in Vision-Language Models (VLMs) that rely on fixed cameras for open-loop tasks such as safety monitoring and operational logging. Evaluating 17 models zero-shot across 3,346 media assets, the benchmark found shortfalls concentrated in event verification, referring expressions, and temporal localization of roughly 9 to 24 points at every model scale, while video question answering stayed within 5.3 points of VideoMME and 2D spatial pointing showed no shortfall against BLINK. No system exceeded 55.7 mIoU on temporal localization or 37.3 SODA_c on dense video captioning, and open-weight models led 2D object localization outright, indicating neither scale nor proprietary access explains the pattern.

by read1 min views1 publishedSep 10, 2026

arXiv:2609.09396v1 Announce Type: new Abstract: As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric consumer video. This overlooks a pervasive class of Physical AI: Infrastructure AI, which relies on fixed cameras for open-loop insights like safety monitoring and operational logging. We introduce VANTAGE-Bench, a benchmark measuring this "Infrastructure AI Gap." It spans three operational domains (Logistics, Transportation, and Smart Spaces), unifies image and video evaluation across semantic, spatial, temporal, and spatio-temporal capabilities, and moves beyond multiple-choice to eight task formulations including dense captioning and spatio-temporal grounding. It adds a single-pass trajectory protocol for Single Object Tracking and, to our knowledge, the first such evaluation on fixed-camera infrastructure video, scored against specialist trackers. Annotation spans three regimes over 3,346 media assets: 3,342 video-task annotations, 4,281 image-grounding annotations, and 27,404 detection boxes. Evaluating 17 models zero-shot, we find the shortfall relative to consumer-centric benchmarks is concentrated, not general. Event verification, referring expressions, and temporal localization fall roughly 9 to 24 points at every model scale, while video question answering stays within 5.3 points of VideoMME and 2D spatial pointing shows no shortfall against BLINK. The temporal pillar is weakest in absolute terms: no system exceeds 55.7 mIoU on temporal localization or 37.3 SODA_c on dense video captioning. On tracking, frontier models come within roughly 5 points of specialist trackers over short horizons but separate as the horizon extends. Open-weight models lead 2D object localization outright, so neither scale nor proprietary access explains the pattern. Data, evaluation harness, and leaderboard: https://vantage-bench.org/

── more in #computer-vision 4 stories · sorted by recency
── more on @vantage-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/vantage-bench-evalua…] indexed:0 read:1min 2026-09-10 ·