cd /news/machine-learning/when-can-test-time-adaptation-help-z… · home topics machine-learning article
[ARTICLE · art-65506] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

When Can Test-Time Adaptation Help Zero-Shot CT Vision-Language Models?

A new study from arXiv finds that test-time adaptation (TTA) can improve zero-shot 3D CT vision-language models only under specific conditions, and introduces CARVE (Cardinality-Aware Retained-View Entropy) as the first TTA method for multi-label adaptation in this setting. The research shows that TTA is conditional on preserving the encoder's depth structure and base representation transfer, with depth reduction alone lowering internal AUROC by more than 0.12. CARVE provides consistent improvements across multi-label, three-class, and binary CT tasks when the base model is already discriminative.

read1 min views2 publishedJul 20, 2026
arXiv:2607.15556v1 Announce Type: new
Abstract: 3D CT vision-language models (VLMs) classify abnormalities from text prompts in a zero-shot manner, enabling cross-institution deployment where labels are scarce and clinical tasks shift faster than supervised models can be retrained. A real CT scan, however, typically contains several co-occurring abnormalities, and the reliability of zero-shot multi-label prediction under distribution shift remains poorly understood. Test-time adaptation (TTA) updates a model on unlabeled target scans without source data or target annotations, yet existing TTA methods target multi-class softmax prediction on natural images or 2D medical segmentation, and none addresses unsupervised multi-label adaptation for zero-shot 3D CT VLMs. We study when TTA helps zero-shot 3D CT VLMs. A controlled diagnostic analysis shows that TTA is conditional: the volumetric input must preserve the encoder's depth structure, and the base representation must transfer to the target cohort, with depth reduction alone lowering internal AUROC by more than 0.12. We then focus on the regime where the base model already separates present from absent abnormalities. We introduce CARVE (Cardinality-Aware Retained-View Entropy), the first TTA method for this setting. CARVE estimates a sample-specific positive-label cardinality $\hat{k}$, optimizes a top-$\hat{k}$ objective to preserve co-occurring abnormalities, and performs memory-efficient multi-view adaptation by scoring weak 3D views without gradients before updating on a retained subset. Across contrastive CT-CLIP and anatomy-aware fVLM, CARVE provides the most consistent improvements across multi-label, three-class, and binary CT tasks when the base model is already discriminative. These results establish multi-label TTA for zero-shot 3D CT VLMs as a distinct problem and CARVE as a cardinality-aware solution.
── more in #machine-learning 4 stories · sorted by recency
── more on @carve 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/when-can-test-time-a…] indexed:0 read:1min 2026-07-20 ·