cd /news/machine-learning/enabling-vision-and-cross-modal-lear… · home topics machine-learning article
[ARTICLE · art-136642] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework

Self-supervised pretraining on 3D CTA scans reduced modality imbalance and improved cross-modal integration for multimodal stroke recurrence prediction, according to an arXiv paper (arXiv:2609.22271v1) whose code is publicly available at github.com/ChristianGappGit/SSL_Pretraining. Two multimodal neural networks were pretrained in a self-supervised manner and fine-tuned under two distinct freezing strategies, outperforming both the prior baseline model and all models trained entirely from scratch. The best-performing Vision Transformer based network overcame unimodal collapse, and synergy analysis found significant interactions between vision and both gender and CHD.

by read1 min views1 publishedSep 22, 2026

arXiv:2609.22271v1 Announce Type: new Abstract: Multimodal stroke recurrence prediction requires effective integration of heterogeneous clinical and imaging data, yet modality imbalance often causes models to over-rely on dominant modalities and underutilize complementary information. While self-supervised pretraining and selective parameter freezing are commonly employed to improve representation learning and fine-tuning stability, their effect on modality contributions and cross-modal behavior in multimodal medical models remains largely unexplored. In this work, we investigate whether image pretraining on 3D CTA scans reduces modality imbalance and improves cross-modal integration for stroke recurrence prediction, a clinically critical task we recently addressed. To this end, two multimodal neural networks are pretrained in a self-supervised manner and subsequently fine-tuned using two distinct freezing strategies. Their performance and modality utilization are compared against both the baseline model from our previous work and models trained entirely from scratch in this study. Our results demonstrate that self-supervised pretraining enables more effective utilization of the multimodal image-tabular dataset, outperforming both the prior baseline and all non-pretrained models. Notably, the best-performing Vision Transformer based neural network successfully overcomes unimodal collapse. Synergy analysis reveals significant interactions between vision and both gender and CHD, suggesting clinically relevant patterns for stroke recurrence. Overall, our findings demonstrate that self-supervised pretraining and strategic fine-tuning support more balanced modality utilization and enable meaningful cross-modal interactions. Code is publicly available at https://github.com/ChristianGappGit/SSL_Pretraining.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/enabling-vision-and-…] indexed:0 read:1min 2026-09-22 ·