{"slug": "traceclip-recovering-local-semantics-from-patch-to-cls-contributions", "title": "TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions", "summary": "Researchers introduce TraceCLIP, a training-free framework that recovers latent patch-level semantic evidence from CLIP's CLS attention output, achieving gains of 1.3 to 4.5 points in average mIoU over prior training-free methods on eight zero-shot semantic segmentation benchmarks without additional training or supervision.", "body_md": "arXiv:2607.26107v1 Announce Type: new\nAbstract: Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions. CLIP provides a strong foundation for these tasks by learning a shared image-text embedding space from large-scale contrastive pre-training. However, its image-level objective aligns text with a CLS-derived global representation, leaving local vision-language correspondence only indirectly constrained. Existing methods either introduce additional supervision, external models, or task-specific adaptation, while training-free approaches mainly recover dense responses from existing patch features without examining where local semantics become most accessible within CLIP. We introduce TraceCLIP, a training-free framework that recovers latent patch-level semantic evidence by isolating the patch-specific terms written into the CLS attention output. TraceCLIP further converts contribution-derived semantic responses into a semantic-geodesic topology gate that calibrates final-layer patch affinity for dense feature reconstruction. Diagnostic experiments show that these contribution features exhibit strong local semantic discrimination and text-conditioned spatial alignment. On eight zero-shot semantic segmentation benchmarks, TraceCLIP achieves gains of 1.3 to 4.5 points in average mIoU over the strongest prior training-free methods across both backbones and background settings, without additional training, external vision foundation models, or region-level supervision. More broadly, these findings suggest that spatially localized semantics may remain accessible within the internal construction of globally aligned representations.", "url": "https://wpnews.pro/news/traceclip-recovering-local-semantics-from-patch-to-cls-contributions", "canonical_source": "https://arxiv.org/abs/2607.26107", "published_at": "2026-07-30 04:00:00+00:00", "updated_at": "2026-07-30 04:33:47.440204+00:00", "lang": "en", "topics": ["computer-vision", "machine-learning", "large-language-models", "artificial-intelligence"], "entities": ["TraceCLIP", "CLIP"], "alternates": {"html": "https://wpnews.pro/news/traceclip-recovering-local-semantics-from-patch-to-cls-contributions", "markdown": "https://wpnews.pro/news/traceclip-recovering-local-semantics-from-patch-to-cls-contributions.md", "text": "https://wpnews.pro/news/traceclip-recovering-local-semantics-from-patch-to-cls-contributions.txt", "jsonld": "https://wpnews.pro/news/traceclip-recovering-local-semantics-from-patch-to-cls-contributions.jsonld"}}