{"slug": "generative-semantic-segmentation-via-an-observable-semantic-image-interface-and", "title": "Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment", "summary": "Researchers introduced Semantic Prism, a generative semantic segmentation framework that achieves 72.07% mean intersection over union (mIoU) on the Cityscapes validation set, outperforming direct-interface decoding by 11.39 mIoU points with 0.41% expected calibration error. The framework uses a diffusion-distilled one-step generator and hierarchical feature alignment to improve pixel-level segmentation accuracy. On BDD100K, a separately trained model reached 62.22% mIoU, and the Cityscapes-trained model achieved 46.89% mIoU on the Adverse Conditions Dataset with Correspondences without target-domain adaptation.", "body_md": "arXiv:2608.11537v1 Announce Type: new\nAbstract: Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rendered image to an intermediate visualization. We present Semantic Prism, a conditional semantic-image generation-and-refinement framework with deterministic inference. A diffusion-distilled one-step generator renders a semantic RGB image; per-pixel distances from the rendered colors to a fixed class-color codebook define an explicit probabilistic interface. Hierarchical Generator Evidence Alignment spatially aligns multi-level generator features and uses a zero-initialized output projection to predict an additive residual in the interface logit space, retaining the image-defined interface as the reference for the final distribution. The interface and refined distributions further enable Contextual Interface--Hierarchy Disagreement (C-IHD), a fixed readout for ranking remaining pixel errors without an auxiliary predictor or additional forward pass. On the 500-image Cityscapes validation set, Semantic Prism achieves 72.07% mean intersection over union, 11.39 mIoU points above direct-interface decoding, with 0.41% expected calibration error. Matched-capacity ablations over three seeds support the benefit of jointly aligned multi-level evidence. A separately trained model attains 62.22% mIoU on BDD100K, while the Cityscapes-trained model reaches 46.89\\% mIoU under source-frozen transfer to the Adverse Conditions Dataset with Correspondences, without target-domain adaptation. Across all three datasets, C-IHD consistently improves the area under the precision--recall curve for pixel-error ranking over maximum softmax probability on the same segmentation predictions; on ACDC, it raises AUPR from 0.6580 to 0.7557.", "url": "https://wpnews.pro/news/generative-semantic-segmentation-via-an-observable-semantic-image-interface-and", "canonical_source": "https://arxiv.org/abs/2608.11537", "published_at": "2026-08-13 04:00:00+00:00", "updated_at": "2026-08-13 04:13:15.692801+00:00", "lang": "en", "topics": ["computer-vision", "generative-ai", "machine-learning"], "entities": ["Semantic Prism", "Cityscapes", "BDD100K", "Adverse Conditions Dataset with Correspondences"], "alternates": {"html": "https://wpnews.pro/news/generative-semantic-segmentation-via-an-observable-semantic-image-interface-and", "markdown": "https://wpnews.pro/news/generative-semantic-segmentation-via-an-observable-semantic-image-interface-and.md", "text": "https://wpnews.pro/news/generative-semantic-segmentation-via-an-observable-semantic-image-interface-and.txt", "jsonld": "https://wpnews.pro/news/generative-semantic-segmentation-via-an-observable-semantic-image-interface-and.jsonld"}}