Breast Ultrasound AI Under Dataset Shift : A Patient-Leakage-Aware Benchmark
| dc.contributor.author | Wang, Lulu | |
| dc.contributor.department | Department of Engineering | |
| dc.date.accessioned | 2026-09-02T10:41:01Z | |
| dc.date.available | 2026-09-02T10:41:01Z | |
| dc.date.issued | 2026-05 | |
| dc.description | Publisher Copyright: © 2026 by the author. | en |
| dc.description.abstract | Background: Artificial intelligence (AI) has shown promise in breast ultrasound image analysis, but most evidence still comes from single-dataset studies. Clinical translation requires evaluation under heterogeneous acquisition and curation conditions. This study presents a patient-leakage-aware, reproducible benchmark for breast ultrasound AI under dataset shift, with emphasis on external generalization, calibration, and confidence-related behavior. Methods: A reproducible benchmark framework was developed using patient-level splitting, internal testing, pairwise cross-dataset evaluation, whole-image and region-of-interest (ROI) input strategies, calibration analysis, targeted ROI-margin sensitivity analysis, representative explainable AI visualization, and an auxiliary lesion-versus-normal confidence-based analysis. Four public breast ultrasound datasets (BUSI, BUS-UCLM, BUS-BRA, and BrEaST) were harmonized for a primary benign-versus-malignant lesion classification task. Normal images were excluded from the primary endpoint and used only in auxiliary analyses when sufficient numbers were available. Results: Cross-dataset testing was weaker on average than internal testing, with mean raw AUROC decreasing from 0.801 to 0.719 and mean balanced accuracy from 0.723 to 0.635. ROI input improved external performance, especially for the vision transformer, increasing mean external AUROC from 0.666 to 0.805 and mean external balanced accuracy from 0.594 to 0.713 relative to whole-image input. Temperature scaling improved calibration-related metrics, reducing mean external expected calibration error from 0.180 to 0.150 and mean external negative log-likelihood from 0.848 to 0.682. Conclusions: This study establishes a reproducible benchmark for evaluating breast ultrasound AI under dataset shift, with explicit attention to patient-level leakage control, external validity, and reliability of predicted probabilities. | en |
| dc.description.version | Peer reviewed | en |
| dc.format.extent | 1669088 | |
| dc.format.extent | ||
| dc.identifier.citation | Wang, L 2026, 'Breast Ultrasound AI Under Dataset Shift : A Patient-Leakage-Aware Benchmark', Diagnostics, vol. 16, no. 10, 1537. https://doi.org/10.3390/diagnostics16101537 | en |
| dc.identifier.doi | 10.3390/diagnostics16101537 | |
| dc.identifier.issn | 2075-4418 | |
| dc.identifier.other | 250695019 | |
| dc.identifier.other | 44945c4a-61c4-4099-a1d9-05a3772703bf | |
| dc.identifier.other | 105040192190 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.11815/8119 | |
| dc.language.iso | en | |
| dc.relation.ispartofseries | Diagnostics; 16(10) | en |
| dc.relation.url | https://www.scopus.com/pages/publications/105040192190 | en |
| dc.rights | info:eu-repo/semantics/openAccess | en |
| dc.subject | artificial intelligence | en |
| dc.subject | breast ultrasound | en |
| dc.subject | calibration | en |
| dc.subject | dataset shift | en |
| dc.subject | diagnostic imaging | en |
| dc.subject | Internal Medicine | en |
| dc.subject | Clinical Biochemistry | en |
| dc.title | Breast Ultrasound AI Under Dataset Shift : A Patient-Leakage-Aware Benchmark | en |
| dc.type | /dk/atira/pure/researchoutput/researchoutputtypes/contributiontojournal/article | en |
Skrár
Original bundle
1 - 1 af 1
- Nafn:
- diagnostics-16-01537-with-cover.pdf
- Stærð:
- 1.59 MB
- Snið:
- Adobe Portable Document Format