{"slug": "large-language-models-as-unified-multimodal-learners-for-clinical-prediction", "title": "Large Language Models as Unified Multimodal Learners for Clinical Prediction", "summary": "A new study from researchers evaluating large language models as unified multimodal learners for clinical prediction finds that converting all patient data into a single natural language sequence and fine-tuning a pretrained language model matches or exceeds task-specific multimodal baselines across three clinical tasks. The approach, tested on in-hospital mortality using MIMIC-III, graft failure prediction from a German transplant center, and emergency triage classification, outperformed a clinically deployed gradient boosting system for graft failure prediction without requiring bespoke fusion architectures.", "body_md": "arXiv:2607.15380v1 Announce Type: new\nAbstract: Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities. Yet most clinical prediction systems still rely on task-specific fusion architectures, pairing dedicated encoders for each modality with learned combination mechanisms that must be re-engineered for every new task and clinical setting. We propose a simpler alternative: convert all patient data, regardless of modality, into a single natural language sequence and fine-tune a pretrained language model end-to-end, with no architectural modification for fusion. We evaluate this approach across three clinically distinct prediction tasks: in-hospital mortality on MIMIC-III, graft failure prediction using longitudinal data from a German transplant center, and emergency triage classification from ambulance records - comparing encoder-based (ModernBERT) and decoder-based (Llama 3.1, Gemma, DeepSeek-R1-Qwen, Qwen3) fine-tuning against established multimodal baselines and, for graft failure, a gradient boosting model currently used in clinical practice for post-transplant patient management. Across all three tasks, unified textual serialization matches or exceeds task-specific multimodal baselines, and outperforms the clinically deployed gradient boosting system on graft failure prediction. These results indicate that a single serialization-based paradigm, without bespoke fusion architectures, is sufficient for multimodal clinical prediction - substantially reducing system complexity while matching or exceeding specialized designs.", "url": "https://wpnews.pro/news/large-language-models-as-unified-multimodal-learners-for-clinical-prediction", "canonical_source": "https://arxiv.org/abs/2607.15380", "published_at": "2026-07-20 04:00:00+00:00", "updated_at": "2026-07-20 13:40:51.372137+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "artificial-intelligence", "ai-research"], "entities": ["MIMIC-III", "ModernBERT", "Llama 3.1", "Gemma", "DeepSeek-R1-Qwen", "Qwen3"], "alternates": {"html": "https://wpnews.pro/news/large-language-models-as-unified-multimodal-learners-for-clinical-prediction", "markdown": "https://wpnews.pro/news/large-language-models-as-unified-multimodal-learners-for-clinical-prediction.md", "text": "https://wpnews.pro/news/large-language-models-as-unified-multimodal-learners-for-clinical-prediction.txt", "jsonld": "https://wpnews.pro/news/large-language-models-as-unified-multimodal-learners-for-clinical-prediction.jsonld"}}