{"slug": "artemis-geometry-grounded-multi-agent-driving-world-models-with-shared-3d-state", "title": "Artemis: Geometry-Grounded Multi-Agent Driving World Models with Shared 3D State and Progressive Memory Update", "summary": "Researchers posted arXiv paper 2610.07031v1, titled \"Artemis: Geometry-Grounded Multi-Agent Driving World Models with Shared 3D State and Progressive Memory Update,\" which proposes a multi-agent driving world model that reconstructs an explicit 3D world map from multi-agent observations to enforce a unified 3D state across agents. Artemis uses an action-guided geometric injection module to render decomposed foreground-background control maps injected into a diffusion transformer via a GeoAdapter block, and progressively updates the reconstructed 3D world maps using keyframes selected from video rollouts. The authors report that experiments on MA-CARLA, a new dataset sampled from the CARLA simulator, show improved visual fidelity and cross-view consistency, with support for simultaneous 2D video and 3D point map rollouts, scaling beyond two agents and multi-camera settings.", "body_md": "arXiv:2610.07031v1 Announce Type: new \nAbstract: Recent video world models have witnessed the paradigm shift from single-agent to multi-agent involvements, which can reveal more complicated dynamics and cross-agent interaction in the real world. However, existing approaches commonly adopt implicit inter-agent communications via cross attention, which lack explicit geometry constraints and unified 3D state, thereby leading to poor multi-view consistency and struggling with recovering out-of-sight agents. In addition, most of them assume a static background, failing to represent uncontrolled background dynamics. To address these problems, we propose Artemis: a geometry-grounded multi-agent world model with explicit memory sharing. An explicit 3D world map is reconstructed from multi-agent observations to enforce a unified 3D state across agents, offering high cross-view consistency. Specifically, an action-guided geometric injection module is developed to simultaneously render decomposed foreground-background control maps, which are then injected into a diffusion transformer through a designed GeoAdapter block. Compared to previous methods assuming static-only background, our GeoAdapter can also distinguish uncontrolled non-agent dynamics, which are conditioned on their own multi-frame history positions to provide consistent motion cues. Keyframes selected from progressive video rollouts are used to progressively update the reconstructed 3D world maps. To effectively capture complex dynamic patterns, we curate a novel dataset sampled from the CARLA simulator called MA-CARLA. Extensive experiments demonstrate the superiority of our proposed method in terms of visual fidelity and cross-view consistency in the generated videos. In addition, our Artemis can support simultaneous multi-modal rollouts with both 2D video and 3D point map maintenance, scale to scenarios beyond two agents and multi-camera setting.", "url": "https://wpnews.pro/news/artemis-geometry-grounded-multi-agent-driving-world-models-with-shared-3d-state", "canonical_source": "https://arxiv.org/abs/2610.07031", "published_at": "2026-10-07 04:00:00+00:00", "updated_at": "2026-10-07 04:16:21.637137+00:00", "lang": "en", "topics": ["autonomous-vehicles", "computer-vision", "ai-research", "generative-ai", "machine-learning"], "entities": ["Artemis", "CARLA", "MA-CARLA", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/artemis-geometry-grounded-multi-agent-driving-world-models-with-shared-3d-state", "markdown": "https://wpnews.pro/news/artemis-geometry-grounded-multi-agent-driving-world-models-with-shared-3d-state.md", "text": "https://wpnews.pro/news/artemis-geometry-grounded-multi-agent-driving-world-models-with-shared-3d-state.txt", "jsonld": "https://wpnews.pro/news/artemis-geometry-grounded-multi-agent-driving-world-models-with-shared-3d-state.jsonld"}}