cd /news/artificial-intelligence/rad-jepa-3d-radiology-joint-embeddin… · home topics artificial-intelligence article
[ARTICLE · art-79709] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography

Rad-JEPA 3D, a joint-embedding predictive framework for 3D CT scans introduced by researchers, achieves state-of-the-art results on closed-ended VQA and the best average spatial-reasoning score on the Spatial-Med benchmark despite having only 4.0B total parameters, after pretraining on approximately 120,000 CT scans. The model uses a hybrid H-Mamba encoder and Hidden States Orthogonal Regularization to improve volumetric representations.

read1 min views1 publishedJul 30, 2026

arXiv:2607.26196v1 Announce Type: new Abstract: Self-supervised pretraining is central to 3D medical image analysis, where unlabeled CT volumes are abundant but expert annotations are scarce. Yet existing volumetric encoders often fail to preserve the coarse spatial and geometric structure that downstream reasoning depends on, limiting their performance on organ disentanglement, abnormality detection, and spatial understanding when paired with language models. We introduce Rad-JEPA 3D, a joint-embedding predictive framework that learns volumetric CT representations by predicting the latent features of a complete scan from a masked view. At its core is a hybrid H-Mamba encoder that fuses a Mamba state-space branch, which models inter-slice continuity through sequential scanning, with a grouped-query attention branch, which captures cross-plane spatial context, combined through a lightweight per-token router. To improve the quality of intermediate representations, we further propose Hidden States Orthogonal Regularization (HSOR), which aligns student-teacher hidden states and reduces feature redundancy throughout the encoder. This layer-wise regularization produces more consistent and discriminative volumetric representations, leading to improved performance on organ recognition and spatial reasoning tasks. Pretrained on approximately 120,000 CT scans, Rad-JEPA 3D attains state-of-the-art results despite its compact size: with only 4.0B total parameters, it achieves competitive results with state-of-the-art on closed-ended VQA and the best average spatial-reasoning score on the Spatial-Med benchmark. Ablation studies confirm that the hybrid block and HSOR contribute complementary gains, and that the induced spatial structure can substitute for raw language-model scale on volumetric reasoning tasks.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @rad-jepa 3d 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/rad-jepa-3d-radiolog…] indexed:0 read:1min 2026-07-30 ·