{"slug": "can-vision-language-models-assess-proxemic-risk-from-egocentric-robot-images", "title": "Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?", "summary": "A study evaluating three open-source vision-language models (VLMs) — InternVL, Qwen-VL, and SmolVLM — on classifying proxemic danger from egocentric robot images found that without fine-tuning, all models perform near a stratified random baseline, while fine-tuning yields only modest overall improvements. However, Qwen-VL with an advanced prompt achieves substantially higher recall for high-danger cases than the other models, and analysis shows correct danger classification does not correspond to better spatial grounding. The findings indicate current VLMs remain limited in fine-grained proxemic reasoning and spatial grounding, though targeted prompting and fine-tuning can improve high-danger detection in selected models.", "body_md": "arXiv:2608.12515v1 Announce Type: new\nAbstract: Assessing proxemic danger from a robot's egocentric perspective is critical for safe embodied navigation in human environments and requires both visual and contextual reasoning. We evaluate three opensource vision-language models (VLMs) (\\textit{InternVL}, \\textit{Qwen-VL}, and \\textit{SmolVLM}) on the classification of egocentric robot images into four danger levels, comparing three prompting strategies and two rounds of QLoRA fine-tuning against a stratified random baseline. Without fine-tuning, all models perform near the baseline, while fine-tuning yields only modest overall improvements. However, \\textit{Qwen-VL} with an advanced prompt achieves substantially higher recall for high-danger cases than the other models. An analysis of person localization further shows that correct danger classification does not correspond to better spatial grounding, indicating that a model may produce a useful safety label without attending to the relevant region of the scene. These results show that current VLMs remain limited in fine-grained proxemic reasoning and spatial grounding, although targeted prompting and fine-tuning can improve high-danger detection in selected models.", "url": "https://wpnews.pro/news/can-vision-language-models-assess-proxemic-risk-from-egocentric-robot-images", "canonical_source": "https://arxiv.org/abs/2608.12515", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 04:06:16.932861+00:00", "lang": "en", "topics": ["computer-vision", "artificial-intelligence", "machine-learning"], "entities": ["InternVL", "Qwen-VL", "SmolVLM", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/can-vision-language-models-assess-proxemic-risk-from-egocentric-robot-images", "markdown": "https://wpnews.pro/news/can-vision-language-models-assess-proxemic-risk-from-egocentric-robot-images.md", "text": "https://wpnews.pro/news/can-vision-language-models-assess-proxemic-risk-from-egocentric-robot-images.txt", "jsonld": "https://wpnews.pro/news/can-vision-language-models-assess-proxemic-risk-from-egocentric-robot-images.jsonld"}}