Breast Ultrasound AI Under Dataset Shift : A Patient-Leakage-Aware Benchmark

dc.contributor.authorWang, Lulu
dc.contributor.departmentDepartment of Engineering
dc.date.accessioned2026-09-02T10:41:01Z
dc.date.available2026-09-02T10:41:01Z
dc.date.issued2026-05
dc.descriptionPublisher Copyright: © 2026 by the author.en
dc.description.abstractBackground: Artificial intelligence (AI) has shown promise in breast ultrasound image analysis, but most evidence still comes from single-dataset studies. Clinical translation requires evaluation under heterogeneous acquisition and curation conditions. This study presents a patient-leakage-aware, reproducible benchmark for breast ultrasound AI under dataset shift, with emphasis on external generalization, calibration, and confidence-related behavior. Methods: A reproducible benchmark framework was developed using patient-level splitting, internal testing, pairwise cross-dataset evaluation, whole-image and region-of-interest (ROI) input strategies, calibration analysis, targeted ROI-margin sensitivity analysis, representative explainable AI visualization, and an auxiliary lesion-versus-normal confidence-based analysis. Four public breast ultrasound datasets (BUSI, BUS-UCLM, BUS-BRA, and BrEaST) were harmonized for a primary benign-versus-malignant lesion classification task. Normal images were excluded from the primary endpoint and used only in auxiliary analyses when sufficient numbers were available. Results: Cross-dataset testing was weaker on average than internal testing, with mean raw AUROC decreasing from 0.801 to 0.719 and mean balanced accuracy from 0.723 to 0.635. ROI input improved external performance, especially for the vision transformer, increasing mean external AUROC from 0.666 to 0.805 and mean external balanced accuracy from 0.594 to 0.713 relative to whole-image input. Temperature scaling improved calibration-related metrics, reducing mean external expected calibration error from 0.180 to 0.150 and mean external negative log-likelihood from 0.848 to 0.682. Conclusions: This study establishes a reproducible benchmark for evaluating breast ultrasound AI under dataset shift, with explicit attention to patient-level leakage control, external validity, and reliability of predicted probabilities.en
dc.description.versionPeer revieweden
dc.format.extent1669088
dc.format.extent
dc.identifier.citationWang, L 2026, 'Breast Ultrasound AI Under Dataset Shift : A Patient-Leakage-Aware Benchmark', Diagnostics, vol. 16, no. 10, 1537. https://doi.org/10.3390/diagnostics16101537en
dc.identifier.doi10.3390/diagnostics16101537
dc.identifier.issn2075-4418
dc.identifier.other250695019
dc.identifier.other44945c4a-61c4-4099-a1d9-05a3772703bf
dc.identifier.other105040192190
dc.identifier.urihttps://hdl.handle.net/20.500.11815/8119
dc.language.isoen
dc.relation.ispartofseriesDiagnostics; 16(10)en
dc.relation.urlhttps://www.scopus.com/pages/publications/105040192190en
dc.rightsinfo:eu-repo/semantics/openAccessen
dc.subjectartificial intelligenceen
dc.subjectbreast ultrasounden
dc.subjectcalibrationen
dc.subjectdataset shiften
dc.subjectdiagnostic imagingen
dc.subjectInternal Medicineen
dc.subjectClinical Biochemistryen
dc.titleBreast Ultrasound AI Under Dataset Shift : A Patient-Leakage-Aware Benchmarken
dc.type/dk/atira/pure/researchoutput/researchoutputtypes/contributiontojournal/articleen

Skrár

Original bundle

Niðurstöður 1 - 1 af 1
Nafn:
diagnostics-16-01537-with-cover.pdf
Stærð:
1.59 MB
Snið:
Adobe Portable Document Format