{"slug": "bridging-language-and-spherical-space-object-centric-control-for-text-to", "title": "Bridging Language and Spherical Space: Object-Centric Control for Text-to-Panorama Generation", "summary": "Researchers propose PanoCtrl, an object-centric framework for text-to-panorama generation that converts textual descriptions into structured object-level spherical conditions, integrating them into the diffusion process. The method introduces PanoParse, a text-conditioned parser predicting object semantics and spherical bounding field-of-view (BFoV) parameters, and PanoControl, which injects object-level guidance into the diffusion transformer. Experiments show PanoCtrl achieves state-of-the-art performance in spatial alignment and image quality.", "body_md": "arXiv:2608.20691v1 Announce Type: new\nAbstract: Panoramic image generation is increasingly important for immersive applications such as virtual reality, augmented reality, and 3D content creation. Unlike perspective images, panoramic images represent a viewer-centered $360^\\circ$ surrounding space, where directional expressions such as left, right, front, and behind play a central role in spatial understanding. However, existing text-to-panorama methods largely rely on implicit spatial reasoning and often fail to faithfully ground object-level directional descriptions in spherical panoramic scenes. A straightforward alternative is to introduce explicit layouts, but requiring manually specified spatial conditions reduces the flexibility of language-based interaction and does not directly resolve the misalignment between egocentric directional language and panoramic image space. To address this issue, we propose PanoCtrl, an object-centric framework for controllable text-to-panorama generation. Our method explicitly bridges natural language and spherical panoramic space by converting textual descriptions into structured object-level spherical conditions and integrating them into the diffusion process. Specifically, we introduce PanoParse, a text-conditioned parser that predicts object semantics and spherical bounding field-of-view (BFoV) parameters, and \\textbf{PanoControl}, which injects object-level semantic and spatial guidance into the diffusion transformer through object-aware attention and spatial residual enhancement. To support this task, we construct PanoGround, a dataset with object-level spherical annotations and diverse directional descriptions for controllable panoramic generation. Extensive experiments demonstrate that PanoCtrl achieves state-of-the-art performance in both spatial alignment and image quality.", "url": "https://wpnews.pro/news/bridging-language-and-spherical-space-object-centric-control-for-text-to", "canonical_source": "https://arxiv.org/abs/2608.20691", "published_at": "2026-08-24 04:00:00+00:00", "updated_at": "2026-08-24 04:16:54.577375+00:00", "lang": "en", "topics": ["generative-ai", "computer-vision", "artificial-intelligence"], "entities": ["PanoCtrl", "PanoParse", "PanoControl", "PanoGround", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/bridging-language-and-spherical-space-object-centric-control-for-text-to", "markdown": "https://wpnews.pro/news/bridging-language-and-spherical-space-object-centric-control-for-text-to.md", "text": "https://wpnews.pro/news/bridging-language-and-spherical-space-object-centric-control-for-text-to.txt", "jsonld": "https://wpnews.pro/news/bridging-language-and-spherical-space-object-centric-control-for-text-to.jsonld"}}