{"slug": "qwen-3d-a-generalist-3d-vision-language-model-for-spatial-understanding", "title": "Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding", "summary": "Alibaba Group's Qwen team introduced Qwen-3D, a geometry-aware large multimodal model that compresses visual information using multi-view geometric cues and 3D Rotary Positional Embeddings, enabling efficient long-horizon spatial reasoning over static scenes. Qwen-3D surpasses existing 3D LMMs and several large proprietary 2D models on diverse benchmarks while maintaining strong 2D vision-language performance through joint training on 2D and 3D data.", "body_md": "arXiv:2608.02980v1 Announce Type: new\nAbstract: Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a natural compression mechanism for visual streams: depth and camera pose enable observations from multiple views and time steps to be fused into a persistent, world-aligned representation. While recent 3D LMMs leverage geometry-aware representations to improve spatial reasoning, they continue to lag behind specialist 3D perception systems on grounding and segmentation tasks. We argue that a key limitation is geometry-aware decoding: existing methods communicate 3D predictions through language tokens, proposal selection, or lightweight grounding queries, creating a bottleneck between language reasoning and dense geometric prediction. Building on these insights, we introduce Qwen-3D, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes. Qwen-3D augments visual tokens with 3D Rotary Positional Embeddings, allowing attention to operate directly in 3D scene space rather than across independent image frames and thereby facilitating scalable cross-view and temporal reasoning. To bridge language and geometry, Qwen-3D incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene representation, unifying referential grounding, instance segmentation, and visual question answering across both images and videos. Across a diverse set of benchmarks, Qwen-3D surpasses existing 3D LMMs and outperforms several large proprietary 2D models. Notably, Qwen-3D achieves these improvements while maintaining strong performance on standard 2D vision-language benchmarks by jointly training on 2D and 3D data.", "url": "https://wpnews.pro/news/qwen-3d-a-generalist-3d-vision-language-model-for-spatial-understanding", "canonical_source": "https://arxiv.org/abs/2608.02980", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 04:05:55.783695+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "large-language-models", "ai-research"], "entities": ["Alibaba Group", "Qwen", "Qwen-3D"], "alternates": {"html": "https://wpnews.pro/news/qwen-3d-a-generalist-3d-vision-language-model-for-spatial-understanding", "markdown": "https://wpnews.pro/news/qwen-3d-a-generalist-3d-vision-language-model-for-spatial-understanding.md", "text": "https://wpnews.pro/news/qwen-3d-a-generalist-3d-vision-language-model-for-spatial-understanding.txt", "jsonld": "https://wpnews.pro/news/qwen-3d-a-generalist-3d-vision-language-model-for-spatial-understanding.jsonld"}}