{"slug": "imagine3d-llm-teaching-mllms-to-imagine-3d-scenes-before-answering", "title": "Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering", "summary": "Researchers introduced Imagine3D-LLM, a multimodal large language model that decodes a small set of learnable summary tokens into a compact 3D Gaussian Splatting representation of a scene, supervised by a photometric reconstruction loss and trained jointly with next-token prediction. The approach, detailed in arXiv paper 2609.38177v1, outperforms prior methods across multiple spatial reasoning and 3D understanding benchmarks, with the authors reporting that reconstruction supervision also strengthens cross-frame correspondence in the model's image features.", "body_md": "arXiv:2609.38177v1 Announce Type: cross \nAbstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.", "url": "https://wpnews.pro/news/imagine3d-llm-teaching-mllms-to-imagine-3d-scenes-before-answering", "canonical_source": "https://www.machinebrief.com/news/imagine3d-llm-teaching-mllms-to-imagine-3d-scenes-before-ans-j8fg", "published_at": "2026-09-30 04:00:00+00:00", "updated_at": "2026-09-30 05:47:06.137373+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "computer-vision", "ai-research"], "entities": ["Imagine3D-LLM", "arXiv", "Multimodal Large Language Models", "3D Gaussian Splatting"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/imagine3d-llm-teaching-mllms-to-imagine-3d-scenes-before-answering", "markdown": "https://wpnews.pro/news/imagine3d-llm-teaching-mllms-to-imagine-3d-scenes-before-answering.md", "text": "https://wpnews.pro/news/imagine3d-llm-teaching-mllms-to-imagine-3d-scenes-before-answering.txt", "jsonld": "https://wpnews.pro/news/imagine3d-llm-teaching-mllms-to-imagine-3d-scenes-before-answering.jsonld"}}