Dataset-Origin Signatures and Shortcut Learning in Screening Mammography AI: A Cross-Dataset Case Study A study from researchers using the Newfoundland and Labrador Breast Screening Dataset (NLBSD) found that supplementing screening mammography AI training data with biopsy-confirmed cases from abnormal-enriched external datasets like CBIS-DDSM and CMMD consistently reduced performance, with the NLBSD-only model achieving an AUC-ROC of 0.737 (95% CI [0.686, 0.785]) versus 0.620–0.644 when external data was added (DeLong test, Holm-corrected p < 0.05). The datasets were separated almost perfectly when the task was reframed as predicting dataset origin, indicating that dataset-specific characteristics introduce domain shift that outweighs the benefit of additional positive cases. arXiv:2607.15416v1 Announce Type: new Abstract: Reliable AI for screening mammography requires training data representative of the low cancer prevalence and subtle abnormalities found in screening populations. We examined whether supplementing such data with biopsy-confirmed cases from abnormal-enriched external datasets improves performance. Using the Newfoundland and Labrador Breast Screening Dataset NLBSD alongside CBIS-DDSM and CMMD, we evaluated an EfficientNet-B5 encoder initialized with Mammo-CLIP weights as a frozen linear probe under consistent preprocessing and patient-level splits. The NLBSD-only model achieved an AUC-ROC of 0.737 95% CI 0.686, 0.785 . Adding external positive cases reduced performance in every configuration AUC-ROC = 0.620--0.644; DeLong test, Holm-corrected $p < 0.05$ , with degradation increasing as additional sources were introduced. Domain-matched evaluation produced modest gains only when the training and test domains coincided, and no configuration surpassed the NLBSD-only model. As a diagnostic, we reframed the task as predicting each examination's dataset of origin. The datasets were separated almost perfectly despite identical preprocessing, indicating that dataset-specific characteristics strongly influence the learned representation. These findings show that na\"ively pooling abnormal-enriched mammography datasets can introduce domain shift that outweighs the benefit of additional positive cases. Differences in acquisition, intensity mapping, and dataset construction persist after normalization, motivating domain-aware strategies for combining heterogeneous mammography datasets.