{"slug": "dataset-origin-signatures-and-shortcut-learning-in-screening-mammography-ai-a", "title": "Dataset-Origin Signatures and Shortcut Learning in Screening Mammography AI: A Cross-Dataset Case Study", "summary": "A study from researchers using the Newfoundland and Labrador Breast Screening Dataset (NLBSD) found that supplementing screening mammography AI training data with biopsy-confirmed cases from abnormal-enriched external datasets like CBIS-DDSM and CMMD consistently reduced performance, with the NLBSD-only model achieving an AUC-ROC of 0.737 (95% CI [0.686, 0.785]) versus 0.620–0.644 when external data was added (DeLong test, Holm-corrected p < 0.05). The datasets were separated almost perfectly when the task was reframed as predicting dataset origin, indicating that dataset-specific characteristics introduce domain shift that outweighs the benefit of additional positive cases.", "body_md": "arXiv:2607.15416v1 Announce Type: new\nAbstract: Reliable AI for screening mammography requires training data representative of the low cancer prevalence and subtle abnormalities found in screening populations. We examined whether supplementing such data with biopsy-confirmed cases from abnormal-enriched external datasets improves performance. Using the Newfoundland and Labrador Breast Screening Dataset (NLBSD) alongside CBIS-DDSM and CMMD, we evaluated an EfficientNet-B5 encoder initialized with Mammo-CLIP weights as a frozen linear probe under consistent preprocessing and patient-level splits.\nThe NLBSD-only model achieved an AUC-ROC of 0.737 (95% CI [0.686, 0.785]). Adding external positive cases reduced performance in every configuration (AUC-ROC = 0.620--0.644; DeLong test, Holm-corrected $p < 0.05$), with degradation increasing as additional sources were introduced. Domain-matched evaluation produced modest gains only when the training and test domains coincided, and no configuration surpassed the NLBSD-only model.\nAs a diagnostic, we reframed the task as predicting each examination's dataset of origin. The datasets were separated almost perfectly despite identical preprocessing, indicating that dataset-specific characteristics strongly influence the learned representation. These findings show that na\\\"ively pooling abnormal-enriched mammography datasets can introduce domain shift that outweighs the benefit of additional positive cases. Differences in acquisition, intensity mapping, and dataset construction persist after normalization, motivating domain-aware strategies for combining heterogeneous mammography datasets.", "url": "https://wpnews.pro/news/dataset-origin-signatures-and-shortcut-learning-in-screening-mammography-ai-a", "canonical_source": "https://arxiv.org/abs/2607.15416", "published_at": "2026-07-20 04:00:00+00:00", "updated_at": "2026-07-20 13:56:23.703407+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research"], "entities": ["Newfoundland and Labrador Breast Screening Dataset", "CBIS-DDSM", "CMMD", "EfficientNet-B5", "Mammo-CLIP"], "alternates": {"html": "https://wpnews.pro/news/dataset-origin-signatures-and-shortcut-learning-in-screening-mammography-ai-a", "markdown": "https://wpnews.pro/news/dataset-origin-signatures-and-shortcut-learning-in-screening-mammography-ai-a.md", "text": "https://wpnews.pro/news/dataset-origin-signatures-and-shortcut-learning-in-screening-mammography-ai-a.txt", "jsonld": "https://wpnews.pro/news/dataset-origin-signatures-and-shortcut-learning-in-screening-mammography-ai-a.jsonld"}}