M$^2$PFN: End-to-End Disentangled Alignment for Generalizable Multimodal In-Context Learning in Alzheimer's Disease Researchers introduced M$^2$PFN, an end-to-end framework that adapts the TabPFN tabular foundation model into a multimodal Alzheimer's disease predictor by back-propagating task gradients into 3D-MRI and tabular encoders while keeping the in-context learning engine frozen. On ADNI (n=2240, three-class CN/MCI/AD), M$^2$PFN reached 65.55% macro-F1 and 82.21% macro-AUC, and by swapping only the head for a TabPFN regressor it predicted baseline MMSE on a 1250-subject sub-cohort with test MAE 1.743. On the external OASIS-3 and SCAN cohorts with no retraining, M$^2$PFN achieved the best AUC and lowest MMSE MAE across all baselines, transferring even when the cognitive instrument changed. arXiv:2609.28836v1 Announce Type: new Abstract: While various multimodal methods combining imaging and tabular data for Alzheimer's disease AD diagnosis were proposed, they are often limited in generalization across cohorts. In-context learning ICL has demonstrated excellent generalization performances and high flexibility in foundational tabular models such as TabPFN. To extend TabPFN's ICL to multimodal AD analysis, the main obstacle is that TabPFN is meta-trained on synthetic tabular priors that do not naturally match the statistical structure of image-derived features. We propose M$^2$PFN, an end-to-end framework that turns this tabular foundation model into a multimodal AD predictor. M$^2$PFN i performs differentiable inference through TabPFN's transformer, back-propagating task gradients into 3D-MRI and tabular encoders; ii aligns the two modalities into a shared subspace, via disentanglement and a contrastive objective, matched to the ICL engine's prior; and iii folds in a frozen tabular-only prediction through a learnable gated shortcut. Because the ICL engine stays frozen, its in-context mechanism is preserved for test-time generalization, while end-to-end training shapes the encoders into features it can exploit. On ADNI $n=2240$, three-class CN/MCI/AD , M$^2$PFN attains $65.55\%$ macro-F1 and $82.21\%$ macro-AUC, surpassing a comprehensive set of unimodal and multimodal baselines. By swapping only the head for a TabPFN regressor, the same architecture regresses baseline MMSE on a $1250$-subject sub-cohort to test MAE $1.743$, outperforming every multimodal baseline. On two external cohorts OASIS-3 and SCAN with no retraining, M$^2$PFN achieves the best AUC and the lowest MMSE MAE across all baselines, and transfers even when the cognitive instrument changes.