{"slug": "covit-instance-correspondence-contrastive-learning-for-vision-transformer", "title": "CoViT: Instance-Correspondence Contrastive Learning for Vision Transformer", "summary": "Researchers propose Contrastive Vision Transformer (CoViT), a self-supervised learning framework that injects instance-awareness into Vision Transformers (ViT) through geometry-guided contrastive learning, achieving stable performance gains of over 2 AP points across multiple instance-level perception tasks without extra decoders or labels. The method coordinates attention maps and embeddings via attention-guided masking and hardest contrastive mining to force ViT to discern subtle differences between object instances.", "body_md": "arXiv:2609.01787v1 Announce Type: new\nAbstract: Vision Transformers (ViT) excel in semantic understanding but fail to discriminate between object instances (e.g., identical embeddings for two dogs), limiting their use in instance-level tasks such as object detection and instance segmentation. We propose Contrastive Vision Transformer (CoViT), a self-supervised learning framework that injects instance-awareness into ViT through geometry-guided contrastive learning. CoViT uniquely coordinates ViT's attention maps and embeddings by constructing triplets: (1) Attention-guided masking: Refine multi-head attention via adaptive thresholding and morphological operations to generate instance masks, identifying foreground anchors; (2) Hardest contrastive mining: For each anchor, computing pairwise embedding similarities to select the intra-instance hardest positive (least similar patch within its mask) and inter-instance hardest negative (most similar patch from other instances), with intra-instance regions masked during negative search. These triplets drive a contrastive loss that simultaneously compresses intra-instance variance and expands inter-instance margins, forcing ViT to discern subtle geometric and appearance differences between instances. CoViT consistently achieves stable performance gains of over 2 AP points across multiple instance-level perception tasks by using ViT as backbone architecture. Notably, CoViT requires no extra decoders or labels, demonstrating that a pure ViT can learn instance-aware representations via inherent attention priors and targeted contrastive constraints. Code and models will be released.", "url": "https://wpnews.pro/news/covit-instance-correspondence-contrastive-learning-for-vision-transformer", "canonical_source": "https://arxiv.org/abs/2609.01787", "published_at": "2026-09-03 04:00:00+00:00", "updated_at": "2026-09-03 04:22:55.651155+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "ai-research"], "entities": ["CoViT", "Vision Transformer (ViT)"], "alternates": {"html": "https://wpnews.pro/news/covit-instance-correspondence-contrastive-learning-for-vision-transformer", "markdown": "https://wpnews.pro/news/covit-instance-correspondence-contrastive-learning-for-vision-transformer.md", "text": "https://wpnews.pro/news/covit-instance-correspondence-contrastive-learning-for-vision-transformer.txt", "jsonld": "https://wpnews.pro/news/covit-instance-correspondence-contrastive-learning-for-vision-transformer.jsonld"}}