Learning Global Camera Poses from Noisy View-Graphs for Structure from Motion A new deep, global Structure-from-Motion framework learns camera poses from noisy view-graphs using a permutation-equivariant, edge-conditioned graph neural network that outputs globally consistent camera extrinsics without ground-truth supervision, relying only on a relative-pose consistency objective. Evaluated on MegaDepth, 1DSfM, Strecha, and BlendedMVS, the method scales to more than a thousand images, achieves superior rotation and translation accuracy versus deep track-centric methods while registering more images across many scenes, and delivers competitive results against state-of-the-art classical pipelines while being much faster. arXiv:2609.09491v1 Announce Type: new Abstract: Camera pose estimation is a key step in 3D reconstruction and view-synthesis pipelines. We present a deep, global Structure-from-Motion framework based on learned view-graph aggregation. Our method employs a permutation-equivariant, edge-conditioned graph neural network that takes noisy pairwise relative poses as input and outputs globally consistent camera extrinsics. The network is trained without ground-truth supervision, relying solely on a relative-pose consistency objective. This is followed by 3D point triangulation and robust bundle adjustment. Our approach is efficient, scalable to more than a thousand images, and robust to graph density. We evaluate our method on MegaDepth, 1DSfM, Strecha, and BlendedMVS. These experiments demonstrate that our method achieves superior rotation and translation accuracy compared to deep track-centric methods while registering more images across many scenes, and competitive results compared to state-of-the-art classical pipelines, while being much faster.