cd /news/artificial-intelligence/mr3d-vl-a-generalist-vision-language… · home topics artificial-intelligence article
[ARTICLE · art-96256] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging

Researchers introduced Mr3D-VL, a 4-billion-parameter vision-language foundation model for multi-parametric 3D magnetic resonance imaging (mpMRI), achieving a BERTScore of 0.856 for report generation, question-answering accuracy of 0.713, and multiple-choice accuracy of 0.912, outperforming existing 4B/7B/30B models. The model employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding to integrate modality and spatial information, addressing limitations in current AI models for brain tumor diagnosis.

read1 min views1 publishedAug 14, 2026

arXiv:2608.12689v1 Announce Type: new Abstract: Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @mr3d-vl 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mr3d-vl-a-generalist…] indexed:0 read:1min 2026-08-14 ·