IDEAS home Printed from https://ideas.repec.org/a/bpj/sagmbi/v16y2017i3p199-216n3.html

Comparing the performance of linear and nonlinear principal components in the context of high-dimensional genomic data integration

Author

Listed:
  • Islam Shofiqul

    (Population Health Research Institute, McMaster University and Hamilton Health Sciences, Hamilton, Ontario, Canada)

  • Anand Sonia

    (Population Health Research Institute, McMaster University and Hamilton Health Sciences, Hamilton, Ontario, Canada)

  • Hamid Jemila

    (Department of Medicine, McMaster University, 1280 Main Street West, Hamilton, Ontario L8S 4K1, Canada)

  • Thabane Lehana

    (Population Health Research Institute, McMaster University and Hamilton Health Sciences, Hamilton, Ontario, Canada)

  • Beyene Joseph

    (Department of Medicine, McMaster University, 1280 Main Street West, Hamilton, Ontario L8S 4K1, Canada)

Abstract

Linear principal component analysis (PCA) is a widely used approach to reduce the dimension of gene or miRNA expression data sets. This method relies on the linearity assumption, which often fails to capture the patterns and relationships inherent in the data. Thus, a nonlinear approach such as kernel PCA might be optimal. We develop a copula-based simulation algorithm that takes into account the degree of dependence and nonlinearity observed in these data sets. Using this algorithm, we conduct an extensive simulation to compare the performance of linear and kernel principal component analysis methods towards data integration and death classification. We also compare these methods using a real data set with gene and miRNA expression of lung cancer patients. First few kernel principal components show poor performance compared to the linear principal components in this occasion. Reducing dimensions using linear PCA and a logistic regression model for classification seems to be adequate for this purpose. Integrating information from multiple data sets using either of these two approaches leads to an improved classification accuracy for the outcome.

Suggested Citation

  • Islam Shofiqul & Anand Sonia & Hamid Jemila & Thabane Lehana & Beyene Joseph, 2017. "Comparing the performance of linear and nonlinear principal components in the context of high-dimensional genomic data integration," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 16(3), pages 199-216, August.
  • Handle: RePEc:bpj:sagmbi:v:16:y:2017:i:3:p:199-216:n:3
    DOI: 10.1515/sagmb-2016-0066
    as

    Download full text from publisher

    File URL: https://doi.org/10.1515/sagmb-2016-0066
    Download Restriction: For access to full text, subscription to the journal or payment for the individual article is required.

    File URL: https://libkey.io/10.1515/sagmb-2016-0066?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to

    for a different version of it.

    References listed on IDEAS

    as
    1. Aguilera, Ana M. & Escabias, Manuel & Valderrama, Mariano J., 2006. "Using principal components for estimating logistic regression with high-dimensional multicollinear data," Computational Statistics & Data Analysis, Elsevier, vol. 50(8), pages 1905-1924, April.
    2. Karatzoglou, Alexandros & Smola, Alexandros & Hornik, Kurt & Zeileis, Achim, 2004. "kernlab - An S4 Package for Kernel Methods in R," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 11(i09).
    3. Jessica Minnier & Ming Yuan & Jun S. Liu & Tianxi Cai, 2015. "Risk Classification With an Adaptive Naive Bayes Kernel Machine Model," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 110(509), pages 393-404, March.
    4. Diana Chang & Alon Keinan, 2014. "Principal Component Analysis Characterizes Shared Pathogenetics from Genome-Wide Association Studies," PLOS Computational Biology, Public Library of Science, vol. 10(9), pages 1-14, September.
    5. Xiaobo Guo & Ye Zhang & Wenhao Hu & Haizhu Tan & Xueqin Wang, 2014. "Inferring Nonlinear Gene Regulatory Networks from Gene Expression Data Based on Distance Correlation," PLOS ONE, Public Library of Science, vol. 9(2), pages 1-7, February.
    6. W. Gibson, 1959. "Three multivariate models: Factor analysis, latent structure analysis, and latent profile analysis," Psychometrika, Springer;The Psychometric Society, vol. 24(3), pages 229-252, September.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Ernest C. Davenport Jr. & Mark L. Davison & Kyungin Park, 2024. "The Use of Reparametrization and Constraints on Linear Models to Parse Qualitative and Quantitative Information for a Set of Predictors," Journal of Educational and Behavioral Statistics, , vol. 49(6), pages 955-975, December.
    2. Tsukioka, Yasutomo & Yanagi, Junya & Takada, Teruko, 2018. "Investor sentiment extracted from internet stock message boards and IPO puzzles," International Review of Economics & Finance, Elsevier, vol. 56(C), pages 205-217.
    3. Batool, Fatima & Hennig, Christian, 2021. "Clustering with the Average Silhouette Width," Computational Statistics & Data Analysis, Elsevier, vol. 158(C).
    4. Daniel J. Luckett & Eric B. Laber & Samer S. El‐Kamary & Cheng Fan & Ravi Jhaveri & Charles M. Perou & Fatma M. Shebl & Michael R. Kosorok, 2021. "Receiver operating characteristic curves and confidence bands for support vector machines," Biometrics, The International Biometric Society, vol. 77(4), pages 1422-1430, December.
    5. Yanzhu Hu & Huiyang Zhao & Xinbo Ai, 2016. "Inferring Weighted Directed Association Network from Multivariate Time Series with a Synthetic Method of Partial Symbolic Transfer Entropy Spectrum and Granger Causality," PLOS ONE, Public Library of Science, vol. 11(11), pages 1-25, November.
    6. van der Linde, Angelika, 2008. "Variational Bayesian functional PCA," Computational Statistics & Data Analysis, Elsevier, vol. 53(2), pages 517-533, December.
    7. Yoon Lee & Sungchul Cho & Haejin Han & Kyoungmin Kim & Yongsuk Hong, 2017. "Heterogeneous Value of Water: Empirical Evidence in South Korea," Sustainability, MDPI, vol. 9(10), pages 1-11, September.
    8. Hermel Homburger & Manuel K Schneider & Sandra Hilfiker & Andreas Lüscher, 2014. "Inferring Behavioral States of Grazing Livestock from High-Frequency Position Data Alone," PLOS ONE, Public Library of Science, vol. 9(12), pages 1-22, December.
    9. Bellotti, Anthony & Brigo, Damiano & Gambetti, Paolo & Vrins, Frédéric, 2021. "Forecasting recovery rates on non-performing loans with machine learning," International Journal of Forecasting, Elsevier, vol. 37(1), pages 428-444.
    10. Riza, Lala Septem & Bergmeir, Christoph & Herrera, Francisco & Benítez, José M., 2015. "frbs: Fuzzy Rule-Based Systems for Classification and Regression in R," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 65(i06).
    11. Karin Wolffhechel & Amanda C Hahn & Hanne Jarmer & Claire I Fisher & Benedict C Jones & Lisa M DeBruine, 2015. "Testing the Utility of a Data-Driven Approach for Assessing BMI from Face Images," PLOS ONE, Public Library of Science, vol. 10(10), pages 1-10, October.
    12. Enrico Toffalini & Paolo Girardi & David Giofrè & Gianmarco Altoè, 2022. "Entia Non Sunt Multiplicanda … Shall I look for clusters in my cognitive data?," PLOS ONE, Public Library of Science, vol. 17(6), pages 1-22, June.
    13. Takahiro Takamatsu & Hideaki Ohtake & Takashi Oozeki, 2022. "Support Vector Quantile Regression for the Post-Processing of Meso-Scale Ensemble Prediction System Data in the Kanto Region: Solar Power Forecast Reducing Overestimation," Energies, MDPI, vol. 15(4), pages 1-18, February.
    14. Lucadamo, Antonio & Camminatiello, Ida & D'Ambra, Antonello, 2021. "A statistical model for evaluating the patient satisfaction," Socio-Economic Planning Sciences, Elsevier, vol. 73(C).
    15. Sara Domínguez-Rodríguez & Miquel Serna-Pascual & Andrea Oletto & Shaun Barnabas & Peter Zuidewind & Els Dobbels & Siva Danaviah & Osee Behuhuma & Maria Grazia Lain & Paula Vaz & Sheila Fernández-Luis, 2022. "Machine learning outperformed logistic regression classification even with limit sample size: A model to predict pediatric HIV mortality and clinical progression to AIDS," PLOS ONE, Public Library of Science, vol. 17(10), pages 1-13, October.
    16. Teruko Takada & Takahiro Kitajima, 2022. "Trend-following with better adaptation to large downside risks," PLOS ONE, Public Library of Science, vol. 17(10), pages 1-31, October.
    17. Sven Husmann & Antoniya Shivarova & Rick Steinert, 2020. "Company classification using machine learning," Papers 2004.01496, arXiv.org, revised May 2020.
    18. Rachel Sippy & Daniel F Farrell & Daniel A Lichtenstein & Ryan Nightingale & Megan A Harris & Joseph Toth & Paris Hantztidiamantis & Nicholas Usher & Cinthya Cueva Aponte & Julio Barzallo Aguilar & An, 2020. "Severity Index for Suspected Arbovirus (SISA): Machine learning for accurate prediction of hospitalization in subjects suspected of arboviral infection," PLOS Neglected Tropical Diseases, Public Library of Science, vol. 14(2), pages 1-20, February.
    19. Nagarajah Varathan & Pushpakanthie Wijekoon, 2019. "Logistic Liu Estimator under stochastic linear restrictions," Statistical Papers, Springer, vol. 60(3), pages 945-962, June.
    20. Hyland, John J. & Heanue, Kevin & McKillop, Jessica & Micha, Evgenia, 2018. "Factors underlying farmers' intentions to adopt best practices: The case of paddock based grazing systems," Agricultural Systems, Elsevier, vol. 162(C), pages 97-106.

    More about this item

    Keywords

    ;
    ;
    ;
    ;
    ;

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:bpj:sagmbi:v:16:y:2017:i:3:p:199-216:n:3. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Peter Golla (email available below). General contact details of provider: https://www.degruyterbrill.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.