{"slug": "vlalight-lightweight-vision-language-action-models-for-emergency-aware-traffic", "title": "VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control", "summary": "Researchers proposed VLALight, a lightweight end-to-end vision-language-action framework that maps intersection observations and signal-phase information directly to discrete traffic signal actions using a compact 0.5 B-parameter model. In experiments, VLALight reduced pooled emergency-vehicle waiting time by 21.1% compared with the cascaded VLMLight method while running in real time on local hardware and generalizing to unseen intersection topologies and traffic-flow patterns. The framework combines multiple directional camera views into a unified visual input and uses textual instructions to link those views to traffic movements and signal phases, avoiding intermediate image-to-text descriptions and handcrafted traffic-state representations.", "body_md": "arXiv:2609.30709v1 Announce Type: new \nAbstract: Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, the loose coupling and repeated information conversion between modules can lead to the loss of fine-grained visual details, while sequential inference introduces substantial latency. To address these limitations, we propose VLALight, a lightweight end-to-end vision-language-action framework that directly maps intersection observations and signal-phase information to discrete signal actions. To handle the multi-view nature of TSC, VLALight combines multiple directional camera views into a unified visual input and uses textual instructions to establish their correspondence with traffic movements and signal phases. This design enables direct action prediction with a compact 0.5 B-parameter model, without intermediate image-to-text descriptions or handcrafted traffic-state representations. Experiments show that VLALight delivers the best emergency-vehicle service of all compared methods, reducing pooled emergency waiting time by 21.1% over the cascaded VLMLight while running in real time on local hardware and generalizing to unseen intersection topologies and traffic-flow patterns.", "url": "https://wpnews.pro/news/vlalight-lightweight-vision-language-action-models-for-emergency-aware-traffic", "canonical_source": "https://arxiv.org/abs/2609.30709", "published_at": "2026-09-28 04:00:00+00:00", "updated_at": "2026-09-28 04:20:56.451200+00:00", "lang": "en", "topics": ["autonomous-vehicles", "computer-vision", "large-language-models", "ai-research", "machine-learning"], "entities": ["VLALight", "VLMLight", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/vlalight-lightweight-vision-language-action-models-for-emergency-aware-traffic", "markdown": "https://wpnews.pro/news/vlalight-lightweight-vision-language-action-models-for-emergency-aware-traffic.md", "text": "https://wpnews.pro/news/vlalight-lightweight-vision-language-action-models-for-emergency-aware-traffic.txt", "jsonld": "https://wpnews.pro/news/vlalight-lightweight-vision-language-action-models-for-emergency-aware-traffic.jsonld"}}