cd /news/computer-vision/one-geometry-different-outcomes-read… · home › topics › computer-vision › article
[ARTICLE · art-142258] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=· neutral

One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

A new arXiv paper (2609.36101v1) reports that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation in CLIP and SigLIP vision-language encoders, explaining why modifying the modality gap helps some tasks and hurts others. The authors decompose the similarity score into three task-specific roles: in zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias; in standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information and induces a multiplicative ranking distortion, with a geometry-derived exponent tracking the grid-search optimum at Spearman rho = 0.93. In mixed-modal retrieval, the gap direction sorts candidates by modality and its removal can improve cross-modal ranking, unlike random or non-gap controls.

by read1 min views1 publishedSep 30, 2026

arXiv:2609.36101v1 Announce Type: new Abstract: Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three task-specific roles. In zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias. In standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information, inducing a multiplicative ranking distortion; a geometry-derived exponent tracks the grid-search optimum (Spearman rho = 0.93) and restores performance in some settings, although the gains transfer unevenly. In mixed-modal retrieval, the gap direction sorts candidates by modality; its removal can improve cross-modal ranking, unlike random or non-gap controls. Residual semantic structure after removal defines the limits of the rank-one account. Together, these results explain why gap modification can improve, degrade, or restore performance across downstream settings. By clarifying when and why gap modification changes model behavior, this account provides a principled basis for selecting gap interventions in similarity-based vision-language systems across evaluated downstream tasks.

── more in #computer-vision 4 stories · sorted by recency
── more on @clip 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/one-geometry-differe…] indexed:0 read:1min 2026-09-30 · —