cd /news/artificial-intelligence/dataset-origin-signatures-and-shortc… · home topics artificial-intelligence article
[ARTICLE · art-65500] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

Dataset-Origin Signatures and Shortcut Learning in Screening Mammography AI: A Cross-Dataset Case Study

Researchers found that adding biopsy-confirmed cases from external datasets to a screening mammography AI model trained on the Newfoundland and Labrador Breast Screening Dataset (NLBSD) reduced performance, with AUC-ROC dropping from 0.737 to as low as 0.620. The degradation occurred because dataset-specific characteristics caused domain shift, which outweighed the benefit of additional positive cases. The study highlights the need for domain-aware strategies when combining heterogeneous mammography datasets.

read1 min views1 publishedJul 20, 2026

arXiv:2607.15416v1 Announce Type: new Abstract: Reliable AI for screening mammography requires training data representative of the low cancer prevalence and subtle abnormalities found in screening populations. We examined whether supplementing such data with biopsy-confirmed cases from abnormal-enriched external datasets improves performance. Using the Newfoundland and Labrador Breast Screening Dataset (NLBSD) alongside CBIS-DDSM and CMMD, we evaluated an EfficientNet-B5 encoder initialized with Mammo-CLIP weights as a frozen linear probe under consistent preprocessing and patient-level splits. The NLBSD-only model achieved an AUC-ROC of 0.737 (95% CI [0.686, 0.785]). Adding external positive cases reduced performance in every configuration (AUC-ROC = 0.620--0.644; DeLong test, Holm-corrected $p < 0.05$), with degradation increasing as additional sources were introduced. Domain-matched evaluation produced modest gains only when the training and test domains coincided, and no configuration surpassed the NLBSD-only model. As a diagnostic, we reframed the task as predicting each examination's dataset of origin. The datasets were separated almost perfectly despite identical preprocessing, indicating that dataset-specific characteristics strongly influence the learned representation. These findings show that na"ively pooling abnormal-enriched mammography datasets can introduce domain shift that outweighs the benefit of additional positive cases. Differences in acquisition, intensity mapping, and dataset construction persist after normalization, motivating domain-aware strategies for combining heterogeneous mammography datasets.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @newfoundland and labrador breast screening dataset 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dataset-origin-signa…] indexed:0 read:1min 2026-07-20 ·