IDEAS home Printed from https://ideas.repec.org/a/plo/pcbi00/1014735.html

Biomarker discovery and patient stratification in pancreatic cancer using incomplete multi-omics data

Author

Listed:
  • Alejandra Paja-García
  • Rafael Romero-Becerra
  • Tero Aittokallio
  • Alberto López

Abstract

Pancreatic ductal adenocarcinoma (PDAC), with a 12% 5-year survival rate, is the most aggressive type of cancer. Early diagnosis for this pathology is rare, and conventional treatments such as surgery, radio- or chemotherapy, have little to no effect on reducing mortality. Machine learning (ML) approaches could be used to identify biomarkers that help clinicians stratify patients and improve treatment outcomes. However, most ML techniques perform poorly with incomplete data, which is usually the case in real-world settings, often forcing researchers to discard valuable information. In this study, unsupervised ML algorithms capable of dealing with missing modalities were applied to incomplete multi-omics data from PDAC patients to identify clinically meaningful patient subgroups. Through a large-scale clustering benchmark including six omics layers, we discovered two novel subgroups with statistically significant differences in survival and recurrence after surgery, particularly within the first two years, when most patient deaths occur, as well as distinct tumor mutational burden. Comprehensive multi-omics analyses revealed substantial molecular differences between patients in both groups, identified three methylation biomarkers to stratify patients, and highlighted dysregulation in key oncogenic pathways. Importantly, the identified groups are different from previous PDAC classifications, both in their patient composition, prognosis, and in the oncogenic gene pathway profiles exhibited. Using an independent cohort, we further demonstrated that both the prognostic value of these subtypes and their underlying biological characteristics are reproducible. These results could lead to better stratified treatment regimens to improve the prognosis of PDAC patients.Author summary: Pancreatic cancer has one of the highest mortality rates among all cancer types, and it is difficult to find common molecular characteristics across tumors that can be used to create new treatments for patients. Machine learning algorithms are been used to analyze the molecular data of cancer patients and identify similar groups, for example, patients that share the same mutations in specific genes. This information can then be used to create tailored therapies for individuals and obtain the best patient outcomes. We applied machine learning algorithms to molecular data from pancreatic cancer patients (e.g.,: DNA, mutations, proteins) and identified two novel groups, with one of the groups showing lower survival probability, and an increased risk of cancer recurrence compared with the other. Further analysis revealed key cellular pathways and that three molecules could be used to separate patients into the two groups. This opens avenues for therapies personalized for every individual, to ensure that pancreatic cancer patients receive the best care they can.

Suggested Citation

  • Alejandra Paja-García & Rafael Romero-Becerra & Tero Aittokallio & Alberto López, 2026. "Biomarker discovery and patient stratification in pancreatic cancer using incomplete multi-omics data," PLOS Computational Biology, Public Library of Science, vol. 22(9), pages 1-29, September.
  • Handle: RePEc:plo:pcbi00:1014735
    DOI: 10.1371/journal.pcbi.1014735
    as

    Download full text from publisher

    File URL: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1014735
    Download Restriction: no

    File URL: https://journals.plos.org/ploscompbiol/article/file?id=10.1371/journal.pcbi.1014735&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pcbi.1014735?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pcbi00:1014735. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: ploscompbiol (email available below). General contact details of provider: https://journals.plos.org/ploscompbiol/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.