IDEAS home Printed from https://ideas.repec.org/a/plo/pdig00/0001706.html

Explainable machine learning for breast cancer prediction in resource-constrained settings: A multi-algorithmic framework integrating shap-based transparency with clinical decision support

Author

Listed:
  • Oluwaseun Adebayo Bamodu
  • Sumaiya Nezam
  • Chen-Chih Chung

Abstract

Breast cancer remains the most commonly diagnosed malignancy among women globally, with disproportionately higher mortality rates in low- and middle-income countries (LMICs) where diagnostic delays and limited specialist pathology capacity are widespread. While machine learning (ML) approaches achieve strong predictive performance for cancer classification, algorithmic opacity and absence of interpretability frameworks tailored to resource-constrained environments have impeded clinical adoption. This study bridges the translational gap between predictive accuracy and clinical utility by developing an explainable artificial intelligence (XAI) framework specifically designed for breast cancer diagnosis in underserved healthcare settings. Using the Wisconsin Breast Cancer Diagnostic Dataset (569 fine-needle aspirate cytological specimens with 30 nuclear morphometric features), we systematically benchmarked eight supervised classification algorithms: Logistic Regression, Random Forest, XGBoost, LightGBM, Support Vector Machine (SVM), Gradient Boosting, Decision Tree, and K-Nearest Neighbors, using stratified 10-fold cross-validation and an independent hold-out test set (80:20 split). Performance was evaluated across discriminative and probabilistic metrics, including AUC-ROC, F1-score, Matthews Correlation Coefficient (MCC), and Brier score, and interpretability was operationalized through SHapley Additive exPlanations (SHAP) analysis with global feature importance, cross-model consensus ranking, and individual-level dependence characterization. All ensemble and regularized models achieved test-set AUCs above 0.98, with XGBoost and SVM attaining the highest AUC of 0.996, and Logistic Regression the highest accuracy (98.25%) and MCC (0.962). SHAP analysis consistently identified worst perimeter, worst concave points, and worst area as the dominant predictors, with strong concordance across gradient-boosted models (pairwise Spearman rho: XGBoost–LightGBM 0.86, XGBoost–Random Forest 0.82, Random Forest–LightGBM 0.67). Logistic Regression also demonstrated superior probability calibration, a critical requirement for clinical risk stratification. Collectively, these findings deliver a reproducible, transparent framework whose SHAP-derived signatures align with established cytopathological principles, supporting responsible integration of interpretable ML into resource-limited diagnostic workflows and providing a template for equitable AI deployment in global oncology.Author summary: Breast cancer is the most commonly diagnosed cancer among women worldwide, yet survival rates remain dramatically lower in low- and middle-income countries (LMICs) compared to high-income nations, largely due to limited access to specialist diagnostic expertise. Machine learning (ML) holds promise for supporting cancer diagnosis in these settings, but most high-performing ML systems function as difficult-to-interpret ‘black boxes,’ undermining clinician trust and adoption, particularly in environments where algorithmic outputs cannot be readily verified by specialist pathologists. In this study, we developed and evaluated an interpretable ML framework for breast cancer classification using fine-needle aspiration cytology data routinely collected in resource-limited settings. By systematically comparing eight ML algorithms and applying SHAP (SHapley Additive exPlanations)-based explainability analysis, we identified which cellular features most strongly influence diagnostic predictions and demonstrated that these features align with established cytopathological criteria. Crucially, we show that simpler, more interpretable models achieve performance comparable to complex ensembles while offering superior probability calibration, a critical property for clinical triage decisions. Our findings support the use of transparent, computationally lightweight ML systems as viable decision-support tools in underserved healthcare environments, advancing the goal of equitable cancer diagnosis globally.

Suggested Citation

  • Oluwaseun Adebayo Bamodu & Sumaiya Nezam & Chen-Chih Chung, 2026. "Explainable machine learning for breast cancer prediction in resource-constrained settings: A multi-algorithmic framework integrating shap-based transparency with clinical decision support," PLOS Digital Health, Public Library of Science, vol. 5(9), pages 1-17, September.
  • Handle: RePEc:plo:pdig00:0001706
    DOI: 10.1371/journal.pdig.0001706
    as

    Download full text from publisher

    File URL: https://journals.plos.org/digitalhealth/article?id=10.1371/journal.pdig.0001706
    Download Restriction: no

    File URL: https://journals.plos.org/digitalhealth/article/file?id=10.1371/journal.pdig.0001706&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pdig.0001706?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pdig00:0001706. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: digitalhealth (email available below). General contact details of provider: https://journals.plos.org/digitalhealth .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.