cd /news/computer-vision/can-vision-language-models-assess-pr… · home topics computer-vision article
[ARTICLE · art-96252] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=· neutral

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?

A study evaluating three open-source vision-language models (VLMs) — InternVL, Qwen-VL, and SmolVLM — on classifying proxemic danger from egocentric robot images found that without fine-tuning, all models perform near a stratified random baseline, while fine-tuning yields only modest overall improvements. However, Qwen-VL with an advanced prompt achieves substantially higher recall for high-danger cases than the other models, and analysis shows correct danger classification does not correspond to better spatial grounding. The findings indicate current VLMs remain limited in fine-grained proxemic reasoning and spatial grounding, though targeted prompting and fine-tuning can improve high-danger detection in selected models.

read1 min views1 publishedAug 14, 2026

arXiv:2608.12515v1 Announce Type: new Abstract: Assessing proxemic danger from a robot's egocentric perspective is critical for safe embodied navigation in human environments and requires both visual and contextual reasoning. We evaluate three opensource vision-language models (VLMs) (\textit{InternVL}, \textit{Qwen-VL}, and \textit{SmolVLM}) on the classification of egocentric robot images into four danger levels, comparing three prompting strategies and two rounds of QLoRA fine-tuning against a stratified random baseline. Without fine-tuning, all models perform near the baseline, while fine-tuning yields only modest overall improvements. However, \textit{Qwen-VL} with an advanced prompt achieves substantially higher recall for high-danger cases than the other models. An analysis of person localization further shows that correct danger classification does not correspond to better spatial grounding, indicating that a model may produce a useful safety label without attending to the relevant region of the scene. These results show that current VLMs remain limited in fine-grained proxemic reasoning and spatial grounding, although targeted prompting and fine-tuning can improve high-danger detection in selected models.

── more in #computer-vision 4 stories · sorted by recency
── more on @internvl 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/can-vision-language-…] indexed:0 read:1min 2026-08-14 ·