cd /news/artificial-intelligence/qwen-3d-a-generalist-3d-vision-langu… · home topics artificial-intelligence article
[ARTICLE · art-87128] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Alibaba Group's Qwen team introduced Qwen-3D, a geometry-aware large multimodal model that compresses visual information using multi-view geometric cues and 3D Rotary Positional Embeddings, enabling efficient long-horizon spatial reasoning over static scenes. Qwen-3D surpasses existing 3D LMMs and several large proprietary 2D models on diverse benchmarks while maintaining strong 2D vision-language performance through joint training on 2D and 3D data.

read1 min views1 publishedAug 5, 2026

arXiv:2608.02980v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a natural compression mechanism for visual streams: depth and camera pose enable observations from multiple views and time steps to be fused into a persistent, world-aligned representation. While recent 3D LMMs leverage geometry-aware representations to improve spatial reasoning, they continue to lag behind specialist 3D perception systems on grounding and segmentation tasks. We argue that a key limitation is geometry-aware decoding: existing methods communicate 3D predictions through language tokens, proposal selection, or lightweight grounding queries, creating a bottleneck between language reasoning and dense geometric prediction. Building on these insights, we introduce Qwen-3D, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes. Qwen-3D augments visual tokens with 3D Rotary Positional Embeddings, allowing attention to operate directly in 3D scene space rather than across independent image frames and thereby facilitating scalable cross-view and temporal reasoning. To bridge language and geometry, Qwen-3D incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene representation, unifying referential grounding, instance segmentation, and visual question answering across both images and videos. Across a diverse set of benchmarks, Qwen-3D surpasses existing 3D LMMs and outperforms several large proprietary 2D models. Notably, Qwen-3D achieves these improvements while maintaining strong performance on standard 2D vision-language benchmarks by jointly training on 2D and 3D data.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @alibaba group 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/qwen-3d-a-generalist…] indexed:0 read:1min 2026-08-05 ·