{"slug": "cmu-drive-and-v2v-vla-cooperative-multi-agent-unified-driving-with-reasoning-and", "title": "CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models", "summary": "Researchers introduced CMU-Drive, a closed-loop end-to-end benchmark for evaluating cooperative autonomous driving with multiple connected autonomous vehicles in safety-critical scenarios, and V2V-VLA, a cooperative Vision-Language-Action model that jointly generates driving actions, future waypoints, language reasoning, and communication policies. The team will publicly release the code, benchmark, and model checkpoint to support open-source research.", "body_md": "arXiv:2608.07621v1 Announce Type: new\nAbstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning. We present Cooperative Multi-agent Unified Driving with Reasoning (CMU-Drive), a closed-loop end-to-end benchmark for evaluating cooperative autonomous driving with multiple connected autonomous vehicles (CAVs) operating in safety-critical driving scenarios with background traffic participants. We further propose Vehicle-to-Vehicle Vision-Language-Action (V2V-VLA), a cooperative VLA model that integrates cooperative driving into a single forward pass by jointly generating driving actions, future waypoints, language reasoning, and communication policies. Experiments on CMU-Drive establish the first benchmark and baseline for cooperative VLA driving and provide a foundation for future research on multi-agent, closed-loop, end-to-end cooperative autonomous driving. Our code, benchmark, and model checkpoint will be publicly released to facilitate open-source research.", "url": "https://wpnews.pro/news/cmu-drive-and-v2v-vla-cooperative-multi-agent-unified-driving-with-reasoning-and", "canonical_source": "https://arxiv.org/abs/2608.07621", "published_at": "2026-08-11 04:00:00+00:00", "updated_at": "2026-08-11 04:16:41.381778+00:00", "lang": "en", "topics": ["autonomous-vehicles", "artificial-intelligence", "machine-learning", "computer-vision", "natural-language-processing"], "entities": ["CMU-Drive", "V2V-VLA", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/cmu-drive-and-v2v-vla-cooperative-multi-agent-unified-driving-with-reasoning-and", "markdown": "https://wpnews.pro/news/cmu-drive-and-v2v-vla-cooperative-multi-agent-unified-driving-with-reasoning-and.md", "text": "https://wpnews.pro/news/cmu-drive-and-v2v-vla-cooperative-multi-agent-unified-driving-with-reasoning-and.txt", "jsonld": "https://wpnews.pro/news/cmu-drive-and-v2v-vla-cooperative-multi-agent-unified-driving-with-reasoning-and.jsonld"}}