cd /news/computer-vision/generative-semantic-segmentation-via… · home topics computer-vision article
[ARTICLE · art-94729] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=· neutral

Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment

Researchers introduced Semantic Prism, a generative semantic segmentation framework that achieves 72.07% mean intersection over union (mIoU) on the Cityscapes validation set, outperforming direct-interface decoding by 11.39 mIoU points with 0.41% expected calibration error. The framework uses a diffusion-distilled one-step generator and hierarchical feature alignment to improve pixel-level segmentation accuracy. On BDD100K, a separately trained model reached 62.22% mIoU, and the Cityscapes-trained model achieved 46.89% mIoU on the Adverse Conditions Dataset with Correspondences without target-domain adaptation.

read1 min views1 publishedAug 13, 2026

arXiv:2608.11537v1 Announce Type: new Abstract: Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rendered image to an intermediate visualization. We present Semantic Prism, a conditional semantic-image generation-and-refinement framework with deterministic inference. A diffusion-distilled one-step generator renders a semantic RGB image; per-pixel distances from the rendered colors to a fixed class-color codebook define an explicit probabilistic interface. Hierarchical Generator Evidence Alignment spatially aligns multi-level generator features and uses a zero-initialized output projection to predict an additive residual in the interface logit space, retaining the image-defined interface as the reference for the final distribution. The interface and refined distributions further enable Contextual Interface--Hierarchy Disagreement (C-IHD), a fixed readout for ranking remaining pixel errors without an auxiliary predictor or additional forward pass. On the 500-image Cityscapes validation set, Semantic Prism achieves 72.07% mean intersection over union, 11.39 mIoU points above direct-interface decoding, with 0.41% expected calibration error. Matched-capacity ablations over three seeds support the benefit of jointly aligned multi-level evidence. A separately trained model attains 62.22% mIoU on BDD100K, while the Cityscapes-trained model reaches 46.89% mIoU under source-frozen transfer to the Adverse Conditions Dataset with Correspondences, without target-domain adaptation. Across all three datasets, C-IHD consistently improves the area under the precision--recall curve for pixel-error ranking over maximum softmax probability on the same segmentation predictions; on ACDC, it raises AUPR from 0.6580 to 0.7557.

── more in #computer-vision 4 stories · sorted by recency
── more on @semantic prism 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/generative-semantic-…] indexed:0 read:1min 2026-08-13 ·