cd /news/large-language-models/large-language-models-as-unified-mul… · home topics large-language-models article
[ARTICLE · art-65389] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Large Language Models as Unified Multimodal Learners for Clinical Prediction

Researchers propose converting all patient data, including free-text clinical narratives and structured measurements, into a single natural language sequence and fine-tuning a pretrained language model for clinical prediction. Across three tasks—in-hospital mortality, graft failure, and emergency triage—this unified approach matches or exceeds task-specific multimodal baselines and outperforms a clinically deployed gradient boosting model for graft failure. The method reduces system complexity without requiring bespoke fusion architectures.

read1 min views1 publishedJul 20, 2026

arXiv:2607.15380v1 Announce Type: new Abstract: Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities. Yet most clinical prediction systems still rely on task-specific fusion architectures, pairing dedicated encoders for each modality with learned combination mechanisms that must be re-engineered for every new task and clinical setting. We propose a simpler alternative: convert all patient data, regardless of modality, into a single natural language sequence and fine-tune a pretrained language model end-to-end, with no architectural modification for fusion. We evaluate this approach across three clinically distinct prediction tasks: in-hospital mortality on MIMIC-III, graft failure prediction using longitudinal data from a German transplant center, and emergency triage classification from ambulance records - comparing encoder-based (ModernBERT) and decoder-based (Llama 3.1, Gemma, DeepSeek-R1-Qwen, Qwen3) fine-tuning against established multimodal baselines and, for graft failure, a gradient boosting model currently used in clinical practice for post-transplant patient management. Across all three tasks, unified textual serialization matches or exceeds task-specific multimodal baselines, and outperforms the clinically deployed gradient boosting system on graft failure prediction. These results indicate that a single serialization-based paradigm, without bespoke fusion architectures, is sufficient for multimodal clinical prediction - substantially reducing system complexity while matching or exceeding specialized designs.

── more in #large-language-models 4 stories · sorted by recency
── more on @mimic-iii 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/large-language-model…] indexed:0 read:1min 2026-07-20 ·