IDEAS home Printed from https://ideas.repec.org/a/plo/pcbi00/1010301.html
   My bibliography  Save this article

Archetypal Analysis for population genetics

Author

Listed:
  • Julia Gimbernat-Mayol
  • Albert Dominguez Mantes
  • Carlos D Bustamante
  • Daniel Mas Montserrat
  • Alexander G Ioannidis

Abstract

The estimation of genetic clusters using genomic data has application from genome-wide association studies (GWAS) to demographic history to polygenic risk scores (PRS) and is expected to play an important role in the analyses of increasingly diverse, large-scale cohorts. However, existing methods are computationally-intensive, prohibitively so in the case of nationwide biobanks. Here we explore Archetypal Analysis as an efficient, unsupervised approach for identifying genetic clusters and for associating individuals with them. Such unsupervised approaches help avoid conflating socially constructed ethnic labels with genetic clusters by eliminating the need for exogenous training labels. We show that Archetypal Analysis yields similar cluster structure to existing unsupervised methods such as ADMIXTURE and provides interpretative advantages. More importantly, we show that since Archetypal Analysis can be used with lower-dimensional representations of genetic data, significant reductions in computational time and memory requirements are possible. When Archetypal Analysis is run in such a fashion, it takes several orders of magnitude less compute time than the current standard, ADMIXTURE. Finally, we demonstrate uses ranging across datasets from humans to canids.Author summary: This work introduces a method that combines the singular value decomposition (SVD) with Archetypal Analysis to perform fast and accurate genetic clustering by first reducing the dimensionality of the space of genomic sequences. Each sequence is described as a convex combination (admixture) of archetypes (cluster representatives) in the reduced dimensional space. We compare this interpretable approach to the widely used genetic clustering algorithm, ADMIXTURE, and show that, without significant degradation in performance, Archetypal Analysis outperforms, offering shorter run times and representational advantages. We include theoretical, qualitative, and quantitative comparisons between both methods.

Suggested Citation

  • Julia Gimbernat-Mayol & Albert Dominguez Mantes & Carlos D Bustamante & Daniel Mas Montserrat & Alexander G Ioannidis, 2022. "Archetypal Analysis for population genetics," PLOS Computational Biology, Public Library of Science, vol. 18(8), pages 1-17, August.
  • Handle: RePEc:plo:pcbi00:1010301
    DOI: 10.1371/journal.pcbi.1010301
    as

    Download full text from publisher

    File URL: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1010301
    Download Restriction: no

    File URL: https://journals.plos.org/ploscompbiol/article/file?id=10.1371/journal.pcbi.1010301&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pcbi.1010301?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. Gil McVean, 2009. "A Genealogical Interpretation of Principal Components Analysis," PLOS Genetics, Public Library of Science, vol. 5(10), pages 1-10, October.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Estavoyer, Maxime & François, Olivier, 2022. "Theoretical analysis of principal components in an umbrella model of intraspecific evolution," Theoretical Population Biology, Elsevier, vol. 148(C), pages 11-21.
    2. Pei-Kuan Cong & Wei-Yang Bai & Jin-Chen Li & Meng-Yuan Yang & Saber Khederzadeh & Si-Rui Gai & Nan Li & Yu-Heng Liu & Shi-Hui Yu & Wei-Wei Zhao & Jun-Quan Liu & Yi Sun & Xiao-Wei Zhu & Pian-Pian Zhao , 2022. "Genomic analyses of 10,376 individuals in the Westlake BioBank for Chinese (WBBC) pilot project," Nature Communications, Nature, vol. 13(1), pages 1-15, December.
    3. Ralph, Peter L., 2019. "An empirical approach to demographic inference with genomic data," Theoretical Population Biology, Elsevier, vol. 127(C), pages 91-101.
    4. Peña-Malavera Andrea & Bruno Cecilia & Fernandez Elmer & Balzarini Monica, 2014. "Comparison of algorithms to infer genetic population structure from unlinked molecular markers," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 13(4), pages 391-402, August.
    5. Xiaojun Mao & Somak Dutta & Raymond K. W. Wong & Dan Nettleton, 2020. "Adjusting for Spatial Effects in Genomic Prediction," Journal of Agricultural, Biological and Environmental Statistics, Springer;The International Biometric Society;American Statistical Association, vol. 25(4), pages 699-718, December.
    6. Bryc, Katarzyna & Bryc, Wlodek & Silverstein, Jack W., 2013. "Separation of the largest eigenvalues in eigenanalysis of genotype data from discrete subpopulations," Theoretical Population Biology, Elsevier, vol. 89(C), pages 34-43.
    7. Oscar Lao & Fan Liu & Andreas Wollstein & Manfred Kayser, 2014. "GAGA: A New Algorithm for Genomic Inference of Geographic Ancestry Reveals Fine Level Population Substructure in Europeans," PLOS Computational Biology, Public Library of Science, vol. 10(2), pages 1-11, February.
    8. Alexander Köhler & Marvin Kahra & Michael Breuß, 2024. "A First Approach to Quantum Logical Shape Classification Framework," Mathematics, MDPI, vol. 12(11), pages 1-21, May.
    9. Marie Louis & Petra Korlević & Milaja Nykänen & Frederick Archer & Simon Berrow & Andrew Brownlow & Eline D. Lorenzen & Joanne O’Brien & Klaas Post & Fernando Racimo & Emer Rogan & Patricia E. Rosel &, 2023. "Ancient dolphin genomes reveal rapid repeated adaptation to coastal waters," Nature Communications, Nature, vol. 14(1), pages 1-13, December.
    10. Thomas L. Schmidt & Nancy M. Endersby-Harshman & Anthony R. J. Rooyen & Michelle Katusele & Rebecca Vinit & Leanne J. Robinson & Moses Laman & Stephan Karl & Ary A. Hoffmann, 2024. "Global, asynchronous partial sweeps at multiple insecticide resistance genes in Aedes mosquitoes," Nature Communications, Nature, vol. 15(1), pages 1-19, December.
    11. Priya Moorjani & Nick Patterson & Joel N Hirschhorn & Alon Keinan & Li Hao & Gil Atzmon & Edward Burns & Harry Ostrer & Alkes L Price & David Reich, 2011. "The History of African Gene Flow into Southern Europeans, Levantines, and Jews," PLOS Genetics, Public Library of Science, vol. 7(4), pages 1-13, April.
    12. Yedael Y Waldman & Arjun Biddanda & Natalie R Davidson & Paul Billing-Ross & Maya Dubrovsky & Christopher L Campbell & Carole Oddoux & Eitan Friedman & Gil Atzmon & Eran Halperin & Harry Ostrer & Alon, 2016. "The Genetics of Bene Israel from India Reveals Both Substantial Jewish and Indian Ancestry," PLOS ONE, Public Library of Science, vol. 11(3), pages 1-28, March.
    13. Wang Chaolong & Szpiech Zachary A & Degnan James H & Jakobsson Mattias & Pemberton Trevor J & Hardy John A & Singleton Andrew B & Rosenberg Noah A, 2010. "Comparing Spatial Maps of Human Population-Genetic Variation Using Procrustes Analysis," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 9(1), pages 1-22, January.
    14. Luca Cornetti & Peter D. Fields & Louis Du Pasquier & Dieter Ebert, 2024. "Long-term balancing selection for pathogen resistance maintains trans-species polymorphisms in a planktonic crustacean," Nature Communications, Nature, vol. 15(1), pages 1-11, December.
    15. Zhaoming Wang & Allan Hildesheim & Sophia S Wang & Rolando Herrero & Paula Gonzalez & Laurie Burdette & Amy Hutchinson & Gilles Thomas & Stephen J Chanock & Kai Yu, 2010. "Genetic Admixture and Population Substructure in Guanacaste Costa Rica," PLOS ONE, Public Library of Science, vol. 5(10), pages 1-10, October.
    16. Mofokeng, Maletsema Alina & Mashingaidze, Kingstone, 2018. "Genetic Differentiation of ARC Soybean [Glycine Max (L.) Merrill] Accessions Based on Agronomic and Nutritional Quality Traits," Agriculture and Food Sciences Research, Asian Online Journal Publishing Group, vol. 5(1), pages 6-22.
    17. Gavin Band & Quang Si Le & Luke Jostins & Matti Pirinen & Katja Kivinen & Muminatou Jallow & Fatoumatta Sisay-Joof & Kalifa Bojang & Margaret Pinder & Giorgio Sirugo & David J Conway & Vysaul Nyirongo, 2013. "Imputation-Based Meta-Analysis of Severe Malaria in Three African Populations," PLOS Genetics, Public Library of Science, vol. 9(5), pages 1-13, May.
    18. Duforet-Frebourg, Nicolas & Slatkin, Montgomery, 2016. "Isolation-by-distance-and-time in a stepping-stone model," Theoretical Population Biology, Elsevier, vol. 108(C), pages 24-35.
    19. Alexander Dilthey & Stephen Leslie & Loukas Moutsianas & Judong Shen & Charles Cox & Matthew R Nelson & Gil McVean, 2013. "Multi-Population Classical HLA Type Imputation," PLOS Computational Biology, Public Library of Science, vol. 9(2), pages 1-13, February.
    20. Buschbom, Jutta, 2018. "Exploring and validating statistical reliability in forensic conservation genetics," Thünen Reports 63, Johann Heinrich von Thünen Institute, Federal Research Institute for Rural Areas, Forestry and Fisheries.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pcbi00:1010301. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: ploscompbiol (email available below). General contact details of provider: https://journals.plos.org/ploscompbiol/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.