IDEAS home Printed from https://ideas.repec.org/a/plo/pone00/0237779.html
   My bibliography  Save this article

Predictive analysis methods for human microbiome data with application to Parkinson’s disease

Author

Listed:
  • Mei Dong
  • Longhai Li
  • Man Chen
  • Anthony Kusalik
  • Wei Xu

Abstract

Microbiome data consists of operational taxonomic unit (OTU) counts characterized by zero-inflation, over-dispersion, and grouping structure among samples. Currently, statistical testing methods are commonly performed to identify OTUs that are associated with a phenotype. The limitations of statistical testing methods include that the validity of p-values/q-values depend sensitively on the correctness of models and that the statistical significance does not necessarily imply predictivity. Predictive analysis using methods such as LASSO is an alternative approach for identifying associated OTUs and for measuring the predictability of the phenotype variable with OTUs and other covariate variables. We investigate three strategies of performing predictive analysis: (1) LASSO: fitting a LASSO multinomial logistic regression model to all OTU counts with specific transformation; (2) screening+GLM: screening OTUs with q-values returned by fitting a GLMM to each OTU, then fitting a GLM model using a subset of selected OTUs; (3) screening+LASSO: fitting a LASSO to a subset of OTUs selected with GLMM. We have conducted empirical studies using three simulation datasets generated using Dirichlet-multinomial models and a real gut microbiome data related to Parkinson’s disease to investigate the performance of the three strategies for predictive analysis. Our simulation studies show that the predictive performance of LASSO with appropriate variable transformation works remarkably well on zero-inflated data. Our results of real data analysis show that Parkinson’s disease can be predicted based on selected OTUs after the binary transformation, age, and sex with high accuracy (Error Rate = 0.199, AUC = 0.872, AUPRC = 0.912). These results provide strong evidences of the relationship between Parkinson’s disease and the gut microbiome.

Suggested Citation

  • Mei Dong & Longhai Li & Man Chen & Anthony Kusalik & Wei Xu, 2020. "Predictive analysis methods for human microbiome data with application to Parkinson’s disease," PLOS ONE, Public Library of Science, vol. 15(8), pages 1-18, August.
  • Handle: RePEc:plo:pone00:0237779
    DOI: 10.1371/journal.pone.0237779
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0237779
    Download Restriction: no

    File URL: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0237779&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pone.0237779?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. Tao Wang & Can Yang & Hongyu Zhao, 2019. "Prediction analysis for microbiome sequencing data," Biometrics, The International Biometric Society, vol. 75(3), pages 875-884, September.
    2. Lizhen Xu & Andrew D Paterson & Williams Turpin & Wei Xu, 2015. "Assessment and Selection of Competing Models for Zero-Inflated Microbiome Data," PLOS ONE, Public Library of Science, vol. 10(7), pages 1-30, July.
    3. Fan Xia & Jun Chen & Wing Kam Fung & Hongzhe Li, 2013. "A Logistic Normal Multinomial Regression Model for Microbiome Compositional Data Analysis," Biometrics, The International Biometric Society, vol. 69(4), pages 1053-1063, December.
    4. Tao Wang & Hongyu Zhao, 2017. "Constructing Predictive Microbial Signatures at Multiple Taxonomic Levels," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 112(519), pages 1022-1031, July.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. G. S. Monti & P. Filzmoser, 2022. "Robust logistic zero-sum regression for microbiome compositional data," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 16(2), pages 301-324, June.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Konstantin Shestopaloff & Mei Dong & Fan Gao & Wei Xu, 2021. "Dcmd: Distance-based classification using mixture distributions on microbiome data," PLOS Computational Biology, Public Library of Science, vol. 17(3), pages 1-18, March.
    2. Mozhaeva, Irina, 2022. "Inequalities in utilization of institutional care among older people in Estonia," Health Policy, Elsevier, vol. 126(7), pages 704-714.
    3. Duo Jiang & Thomas Sharpton & Yuan Jiang, 2021. "Microbial Interaction Network Estimation via Bias-Corrected Graphical Lasso," Statistics in Biosciences, Springer;International Chinese Statistical Association, vol. 13(2), pages 329-350, July.
    4. Li, Junlan & Wang, Tao, 2021. "Dimension reduction in binary response regression: A joint modeling approach," Computational Statistics & Data Analysis, Elsevier, vol. 156(C).
    5. Can Cui & Susheela P. Singh & Ana‐Maria Staicu & Brian J. Reich, 2021. "Bayesian variable selection for high‐dimensional rank data," Environmetrics, John Wiley & Sons, Ltd., vol. 32(7), November.
    6. Srinivasan, Arun & Xue, Lingzhou & Zhan, Xiang, 2023. "Identification of microbial features in multivariate regression under false discovery rate control," Computational Statistics & Data Analysis, Elsevier, vol. 181(C).
    7. Pratheepa Jeganathan & Susan P. Holmes, 2021. "A Statistical Perspective on the Challenges in Molecular Microbial Biology," Journal of Agricultural, Biological and Environmental Statistics, Springer;The International Biometric Society;American Statistical Association, vol. 26(2), pages 131-160, June.
    8. Bo Chen & Wei Xu, 2020. "Generalized estimating equation modeling on correlated microbiome sequencing data with longitudinal measures," PLOS Computational Biology, Public Library of Science, vol. 16(9), pages 1-22, September.
    9. Bingkai Wang & Brian S. Caffo & Xi Luo & Chin‐Fu Liu & Andreia V. Faria & Michael I. Miller & Yi Zhao & for the Alzheimer's Disease Neuroimaging Initiative*, 2022. "Regularized regression on compositional trees with application to MRI analysis," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 71(3), pages 541-561, June.
    10. Haixiang Zhang & Jun Chen & Zhigang Li & Lei Liu, 2021. "Testing for Mediation Effect with Application to Human Microbiome Data," Statistics in Biosciences, Springer;International Chinese Statistical Association, vol. 13(2), pages 313-328, July.
    11. Ying Jiang & Linghan Zhang & Junyi Zhang, 2019. "Energy consumption by rural migrant workers and urban residents with a hukou in China: quality-of-life-related factors and built environment," Natural Hazards: Journal of the International Society for the Prevention and Mitigation of Natural Hazards, Springer;International Society for the Prevention and Mitigation of Natural Hazards, vol. 99(3), pages 1431-1453, December.
    12. Peyhardi, Jean & Fernique, Pierre & Durand, Jean-Baptiste, 2021. "Splitting models for multivariate count data," Journal of Multivariate Analysis, Elsevier, vol. 181(C).
    13. Dongyang Yang & Wei Xu, 2023. "Estimation of Mediation Effect on Zero-Inflated Microbiome Mediators," Mathematics, MDPI, vol. 11(13), pages 1-16, June.
    14. Tianchen Xu & Ryan T. Demmer & Gen Li, 2021. "Zero‐inflated Poisson factor model with application to microbiome read counts," Biometrics, The International Biometric Society, vol. 77(1), pages 91-101, March.
    15. Costantino, Francesco & Di Gravio, Giulio & Patriarca, Riccardo & Petrella, Lea, 2018. "Spare parts management for irregular demand items," Omega, Elsevier, vol. 81(C), pages 57-66.
    16. Cindy Xin Feng, 2021. "A comparison of zero-inflated and hurdle models for modeling zero-inflated count data," Journal of Statistical Distributions and Applications, Springer, vol. 8(1), pages 1-19, December.
    17. Tao Wang & Hongyu Zhao, 2017. "A Dirichlet-tree multinomial regression model for associating dietary nutrients with gut microorganisms," Biometrics, The International Biometric Society, vol. 73(3), pages 792-801, September.
    18. Tyler A Joseph & Liat Shenhav & Joao B Xavier & Eran Halperin & Itsik Pe’er, 2020. "Compositional Lotka-Volterra describes microbial dynamics in the simplex," PLOS Computational Biology, Public Library of Science, vol. 16(5), pages 1-22, May.
    19. Zhigang Li & Katherine Lee & Margaret R. Karagas & Juliette C. Madan & Anne G. Hoen & A. James O’Malley & Hongzhe Li, 2018. "Conditional Regression Based on a Multivariate Zero-Inflated Logistic-Normal Model for Microbiome Relative Abundance Data," Statistics in Biosciences, Springer;International Chinese Statistical Association, vol. 10(3), pages 587-608, December.
    20. Yi Zhao & Bingkai Wang & Chin‐Fu Liu & Andreia V. Faria & Michael I. Miller & Brian S. Caffo & Xi Luo, 2023. "Identifying brain hierarchical structures associated with Alzheimer's disease using a regularized regression method with tree predictors," Biometrics, The International Biometric Society, vol. 79(3), pages 2333-2345, September.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0237779. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.