{"slug": "vantage-bench-evaluating-the-infrastructure-ai-gap-in-vision-language-models", "title": "VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models", "summary": "Researchers introduced VANTAGE-Bench, a benchmark measuring an \"Infrastructure AI Gap\" in Vision-Language Models (VLMs) that rely on fixed cameras for open-loop tasks such as safety monitoring and operational logging. Evaluating 17 models zero-shot across 3,346 media assets, the benchmark found shortfalls concentrated in event verification, referring expressions, and temporal localization of roughly 9 to 24 points at every model scale, while video question answering stayed within 5.3 points of VideoMME and 2D spatial pointing showed no shortfall against BLINK. No system exceeded 55.7 mIoU on temporal localization or 37.3 SODA_c on dense video captioning, and open-weight models led 2D object localization outright, indicating neither scale nor proprietary access explains the pattern.", "body_md": "arXiv:2609.09396v1 Announce Type: new \nAbstract: As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric consumer video. This overlooks a pervasive class of Physical AI: Infrastructure AI, which relies on fixed cameras for open-loop insights like safety monitoring and operational logging. We introduce VANTAGE-Bench, a benchmark measuring this \"Infrastructure AI Gap.\" It spans three operational domains (Logistics, Transportation, and Smart Spaces), unifies image and video evaluation across semantic, spatial, temporal, and spatio-temporal capabilities, and moves beyond multiple-choice to eight task formulations including dense captioning and spatio-temporal grounding. It adds a single-pass trajectory protocol for Single Object Tracking and, to our knowledge, the first such evaluation on fixed-camera infrastructure video, scored against specialist trackers. Annotation spans three regimes over 3,346 media assets: 3,342 video-task annotations, 4,281 image-grounding annotations, and 27,404 detection boxes.\n  Evaluating 17 models zero-shot, we find the shortfall relative to consumer-centric benchmarks is concentrated, not general. Event verification, referring expressions, and temporal localization fall roughly 9 to 24 points at every model scale, while video question answering stays within 5.3 points of VideoMME and 2D spatial pointing shows no shortfall against BLINK. The temporal pillar is weakest in absolute terms: no system exceeds 55.7 mIoU on temporal localization or 37.3 SODA_c on dense video captioning. On tracking, frontier models come within roughly 5 points of specialist trackers over short horizons but separate as the horizon extends. Open-weight models lead 2D object localization outright, so neither scale nor proprietary access explains the pattern. Data, evaluation harness, and leaderboard: https://vantage-bench.org/", "url": "https://wpnews.pro/news/vantage-bench-evaluating-the-infrastructure-ai-gap-in-vision-language-models", "canonical_source": "https://arxiv.org/abs/2609.09396", "published_at": "2026-09-10 04:00:00+00:00", "updated_at": "2026-09-10 04:24:33.616122+00:00", "lang": "en", "topics": ["computer-vision", "ai-research", "machine-learning", "large-language-models", "ai-safety"], "entities": ["VANTAGE-Bench", "Vision-Language Models", "VideoMME", "BLINK", "Infrastructure AI", "SODA_c"], "alternates": {"html": "https://wpnews.pro/news/vantage-bench-evaluating-the-infrastructure-ai-gap-in-vision-language-models", "markdown": "https://wpnews.pro/news/vantage-bench-evaluating-the-infrastructure-ai-gap-in-vision-language-models.md", "text": "https://wpnews.pro/news/vantage-bench-evaluating-the-infrastructure-ai-gap-in-vision-language-models.txt", "jsonld": "https://wpnews.pro/news/vantage-bench-evaluating-the-infrastructure-ai-gap-in-vision-language-models.jsonld"}}