Bridging Language and Spherical Space: Object-Centric Control for Text-to-Panorama Generation Researchers propose PanoCtrl, an object-centric framework for text-to-panorama generation that converts textual descriptions into structured object-level spherical conditions, integrating them into the diffusion process. The method introduces PanoParse, a text-conditioned parser predicting object semantics and spherical bounding field-of-view (BFoV) parameters, and PanoControl, which injects object-level guidance into the diffusion transformer. Experiments show PanoCtrl achieves state-of-the-art performance in spatial alignment and image quality. arXiv:2608.20691v1 Announce Type: new Abstract: Panoramic image generation is increasingly important for immersive applications such as virtual reality, augmented reality, and 3D content creation. Unlike perspective images, panoramic images represent a viewer-centered $360^\circ$ surrounding space, where directional expressions such as left, right, front, and behind play a central role in spatial understanding. However, existing text-to-panorama methods largely rely on implicit spatial reasoning and often fail to faithfully ground object-level directional descriptions in spherical panoramic scenes. A straightforward alternative is to introduce explicit layouts, but requiring manually specified spatial conditions reduces the flexibility of language-based interaction and does not directly resolve the misalignment between egocentric directional language and panoramic image space. To address this issue, we propose PanoCtrl, an object-centric framework for controllable text-to-panorama generation. Our method explicitly bridges natural language and spherical panoramic space by converting textual descriptions into structured object-level spherical conditions and integrating them into the diffusion process. Specifically, we introduce PanoParse, a text-conditioned parser that predicts object semantics and spherical bounding field-of-view BFoV parameters, and \textbf{PanoControl}, which injects object-level semantic and spatial guidance into the diffusion transformer through object-aware attention and spatial residual enhancement. To support this task, we construct PanoGround, a dataset with object-level spherical annotations and diverse directional descriptions for controllable panoramic generation. Extensive experiments demonstrate that PanoCtrl achieves state-of-the-art performance in both spatial alignment and image quality.