{"slug": "textslip-text-self-supervised-clip-for-medical-report-generation", "title": "TextSLIP: Text Self-Supervised CLIP for Medical Report Generation", "summary": "Researchers propose TextSLIP, a medical vision-language pretraining framework that augments CLIP with intra-modal text contrastive learning to improve fine-grained semantic supervision for radiology report generation. Pretrained on 7 million brain MRI image-text pairs, TextSLIP outperforms CLIP-style baselines on report generation metrics in controlled comparisons. Ablation studies confirm that text-side self-supervision drives the observed gains, though broader validation across medical domains is needed.", "body_md": "arXiv:2607.21970v1 Announce Type: new\nAbstract: Automating radiology report generation is important for improving reporting consistency and clinical workflows . While Contrastive Language--Image Pretraining (CLIP) has advanced medical vision language modeling, existing CLIP-style approaches may still provide insufficient fine-grained semantic supervision for complex report generation. Standard CLIP primarily optimizes cross-modal alignment, without explicitly structuring the textual embedding space that guides visual representation learning. To address this limitation, we propose TextSLIP, a general medical vision-language pretraining framework that augments CLIP with intra-modal text contrastive learning. By improving textual embedding discriminability through self-supervised augmented text pairs, TextSLIP is designed to provide finer-grained linguistic supervision to the visual encoder. As an initial validation, we pretrain TextSLIP on a curated dataset of 7 million brain MRI image-text pairs and fine-tune the pretrained visual encoder within a report generation architecture. In controlled comparisons with CLIP-style baselines, TextSLIP shows consistent improvements on report generation metrics. Ablation studies further suggest that text-side self-supervision contributes to the observed gains. These results indicate that text-level contrastive learning is a promising direction for improving medical visual-textual alignment, while broader validation across additional medical domains remains an important next step.", "url": "https://wpnews.pro/news/textslip-text-self-supervised-clip-for-medical-report-generation", "canonical_source": "https://arxiv.org/abs/2607.21970", "published_at": "2026-07-27 04:00:00+00:00", "updated_at": "2026-07-27 04:27:59.306519+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "computer-vision", "ai-research"], "entities": ["TextSLIP", "CLIP", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/textslip-text-self-supervised-clip-for-medical-report-generation", "markdown": "https://wpnews.pro/news/textslip-text-self-supervised-clip-for-medical-report-generation.md", "text": "https://wpnews.pro/news/textslip-text-self-supervised-clip-for-medical-report-generation.txt", "jsonld": "https://wpnews.pro/news/textslip-text-self-supervised-clip-for-medical-report-generation.jsonld"}}