cd /news/artificial-intelligence/from-inference-to-adaptation-a-unifi… · home topics artificial-intelligence article
[ARTICLE · art-103914] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

A new arXiv paper (2608.18339v1) proposes a test-time adaptation method for vision-language models (VLMs) that unifies inference and adaptation objectives via optimal transport, improving accuracy by up to 7% over state-of-the-art methods. The method, called \algname, uses a Wasserstein OT formulation to generate robust pseudo-labels and a soft-label InfoNCE loss for adaptation, addressing distribution shifts that cause noisy pseudo-labels.

read1 min views6 publishedAug 20, 2026

arXiv:2608.18339v1 Announce Type: new Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference. Although significant efforts are devoted to adapting VLMs at test time, they rely heavily on noisy pseudo-labels predicted directly from raw embedding similarities during inference, which are unreliable under distribution shift and mislead the adaptation. To avoid noise amplification, existing works craft coarse-grained surrogate objectives during adaptation, which fail to explicitly model sample-level relationships across different modalities, creating objective mismatch with inference, thus leading to marginal performance improvement. In this work, we aim to bridge the detached objectives of inference and adaptation for VLMs, and propose a principled VLM TTA method called \algname. For VLM inference, we formulate the zero-shot image classification task as a cross-modal alignment problem encoded via a Wasserstein OT formulation, providing robust pseudo-labels at the sample-level to effectively adapt VLMs. For VLM adaptation, we adopt a soft-label InfoNCE loss to adapt VLMs based on the OT-induced pseudo-labels, leveraging fine-grained supervisions to explicitly model relationships of individual image-text pairs via contrastive learning, which empowers accurate inference at the same granularity. Moreover, we theoretically reveal that the InfoNCE loss can be neatly reformulated as a Wasserstein OT formulation, thereby unifying the objectives of the inference and adaptation of VLMs to achieve their mutual benefits. Extensive experiments demonstrate the effectiveness and efficiency of our methods, outperforming the best-performing methods by up to 7% with state-of-the-art efficiency.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/from-inference-to-ad…] indexed:0 read:1min 2026-08-20 ·