cd /news/artificial-intelligence/covit-instance-correspondence-contra… · home topics artificial-intelligence article
[ARTICLE · art-119759] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

CoViT: Instance-Correspondence Contrastive Learning for Vision Transformer

Researchers propose Contrastive Vision Transformer (CoViT), a self-supervised learning framework that injects instance-awareness into Vision Transformers (ViT) through geometry-guided contrastive learning, achieving stable performance gains of over 2 AP points across multiple instance-level perception tasks without extra decoders or labels. The method coordinates attention maps and embeddings via attention-guided masking and hardest contrastive mining to force ViT to discern subtle differences between object instances.

read1 min views6 publishedSep 3, 2026

arXiv:2609.01787v1 Announce Type: new Abstract: Vision Transformers (ViT) excel in semantic understanding but fail to discriminate between object instances (e.g., identical embeddings for two dogs), limiting their use in instance-level tasks such as object detection and instance segmentation. We propose Contrastive Vision Transformer (CoViT), a self-supervised learning framework that injects instance-awareness into ViT through geometry-guided contrastive learning. CoViT uniquely coordinates ViT's attention maps and embeddings by constructing triplets: (1) Attention-guided masking: Refine multi-head attention via adaptive thresholding and morphological operations to generate instance masks, identifying foreground anchors; (2) Hardest contrastive mining: For each anchor, computing pairwise embedding similarities to select the intra-instance hardest positive (least similar patch within its mask) and inter-instance hardest negative (most similar patch from other instances), with intra-instance regions masked during negative search. These triplets drive a contrastive loss that simultaneously compresses intra-instance variance and expands inter-instance margins, forcing ViT to discern subtle geometric and appearance differences between instances. CoViT consistently achieves stable performance gains of over 2 AP points across multiple instance-level perception tasks by using ViT as backbone architecture. Notably, CoViT requires no extra decoders or labels, demonstrating that a pure ViT can learn instance-aware representations via inherent attention priors and targeted contrastive constraints. Code and models will be released.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @covit 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/covit-instance-corre…] indexed:0 read:1min 2026-09-03 ·