PhyMo: A Physical-Field Modality for Multimodal AI4Physics Researchers introduced PhyMo, a physics-grounded multimodal framework that organizes heterogeneous measurements through PDE-associated operators, according to an arXiv paper (arXiv:2609.27554v1). PhyMo uses a three-stage procedure: pretraining a physical-field encoder via field reconstruction under PDE residual supervision, aligning its representations with visual embeddings in a shared latent space, and processing fused multimodal representations with downstream prediction heads. Across five datasets spanning diverse physical environments, PhyMo achieved state-of-the-art performance against the strongest baseline on each dataset. arXiv:2609.27554v1 Announce Type: new Abstract: Multimodal learning is emerging as a powerful paradigm for AI for Physics AI4Physics , where predicting physical systems requires the joint interpretation of heterogeneous observations, measurements, and domain knowledge. However, existing approaches typically represent physical quantities and governing equations as generic numerical or textual tokens, overlooking the physical constraints that determine their spatiotemporal interactions. To address this limitation, we introduce the \textbf{physical-field modality} and propose \textbf{PhyMo}, a physics-grounded multimodal framework that organizes heterogeneous measurements through PDE-associated operators. PhyMo follows a three-stage learning procedure: the physical-field encoder is first pretrained through field reconstruction under PDE residual supervision, its representations are subsequently aligned with visual embeddings in a shared latent space, and the fused multimodal representations are finally processed by corresponding downstream prediction heads. Experiments on five datasets spanning diverse physical environments show that PhyMo achieves state-of-the-art performance, compared to the strongest baseline on each dataset, demonstrating the superiority of PhyMo on multimodal representation learning in AI4Physics.