cd /news/computer-vision/traceclip-recovering-local-semantics… · home topics computer-vision article
[ARTICLE · art-79707] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions

Researchers introduce TraceCLIP, a training-free framework that recovers latent patch-level semantic evidence from CLIP's CLS attention output, achieving gains of 1.3 to 4.5 points in average mIoU over prior training-free methods on eight zero-shot semantic segmentation benchmarks without additional training or supervision.

read1 min views1 publishedJul 30, 2026

arXiv:2607.26107v1 Announce Type: new Abstract: Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions. CLIP provides a strong foundation for these tasks by learning a shared image-text embedding space from large-scale contrastive pre-training. However, its image-level objective aligns text with a CLS-derived global representation, leaving local vision-language correspondence only indirectly constrained. Existing methods either introduce additional supervision, external models, or task-specific adaptation, while training-free approaches mainly recover dense responses from existing patch features without examining where local semantics become most accessible within CLIP. We introduce TraceCLIP, a training-free framework that recovers latent patch-level semantic evidence by isolating the patch-specific terms written into the CLS attention output. TraceCLIP further converts contribution-derived semantic responses into a semantic-geodesic topology gate that calibrates final-layer patch affinity for dense feature reconstruction. Diagnostic experiments show that these contribution features exhibit strong local semantic discrimination and text-conditioned spatial alignment. On eight zero-shot semantic segmentation benchmarks, TraceCLIP achieves gains of 1.3 to 4.5 points in average mIoU over the strongest prior training-free methods across both backbones and background settings, without additional training, external vision foundation models, or region-level supervision. More broadly, these findings suggest that spatially localized semantics may remain accessible within the internal construction of globally aligned representations.

── more in #computer-vision 4 stories · sorted by recency
── more on @traceclip 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/traceclip-recovering…] indexed:0 read:1min 2026-07-30 ·