IDEAS home Printed from https://ideas.repec.org/a/plo/pgen00/0020190.html

Population Structure and Eigenanalysis

Author

Listed:
  • Nick Patterson
  • Alkes L Price
  • David Reich

Abstract

Current methods for inferring population structure from genetic data do not provide formal significance tests for population differentiation. We discuss an approach to studying population structure (principal components analysis) that was first applied to genetic data by Cavalli-Sforza and colleagues. We place the method on a solid statistical footing, using results from modern statistics to develop formal significance tests. We also uncover a general “phase change” phenomenon about the ability to detect structure in genetic data, which emerges from the statistical theory we use, and has an important implication for the ability to discover structure in genetic data: for a fixed but large dataset size, divergence between two populations (as measured, for example, by a statistic like FST) below a threshold is essentially undetectable, but a little above threshold, detection will be easy. This means that we can predict the dataset size needed to detect structure.Synopsis: When analyzing genetic data, one often wishes to determine if the samples are from a population that has structure. Can the samples be regarded as randomly chosen from a homogeneous population, or does the data imply that the population is not genetically homogeneous? Patterson, Price, and Reich show that an old method (principal components) together with modern statistics (Tracy–Widom theory) can be combined to yield a fast and effective answer to this question. The technique is simple and practical on the largest datasets, and can be applied both to genetic markers that are biallelic and to markers that are highly polymorphic such as microsatellites. The theory also allows the authors to estimate the data size needed to detect structure if their samples are in fact from two populations that have a given, but small level of differentiation.

Suggested Citation

  • Nick Patterson & Alkes L Price & David Reich, 2006. "Population Structure and Eigenanalysis," PLOS Genetics, Public Library of Science, vol. 2(12), pages 1-20, December.
  • Handle: RePEc:plo:pgen00:0020190
    DOI: 10.1371/journal.pgen.0020190
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.0020190
    Download Restriction: no

    File URL: https://journals.plos.org/plosgenetics/article/file?id=10.1371/journal.pgen.0020190&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pgen.0020190?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. George Nicholson & Albert V. Smith & Frosti Jónsson & Ómar Gústafsson & Kári Stefánsson & Peter Donnelly, 2002. "Assessing population differentiation and isolation from single‐nucleotide polymorphism data," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 64(4), pages 695-715, October.
    2. Noah A Rosenberg & Saurabh Mahajan & Sohini Ramachandran & Chengfeng Zhao & Jonathan K Pritchard & Marcus W Feldman, 2005. "Clines, Clusters, and the Effect of Study Design on the Inference of Human Population Structure," PLOS Genetics, Public Library of Science, vol. 1(6), pages 1-12, December.
    3. Baik, Jinho & Silverstein, Jack W., 2006. "Eigenvalues of large sample covariance matrices of spiked population models," Journal of Multivariate Analysis, Elsevier, vol. 97(6), pages 1382-1408, July.
    4. B. Devlin & Kathryn Roeder, 1999. "Genomic Control for Association Studies," Biometrics, The International Biometric Society, vol. 55(4), pages 997-1004, December.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Peristera Paschou & Petros Drineas & Jamey Lewis & Caroline M Nievergelt & Deborah A Nickerson & Joshua D Smith & Paul M Ridker & Daniel I Chasman & Ronald M Krauss & Elad Ziv, 2008. "Tracing Sub-Structure in the European American Population with PCA-Informative Markers," PLOS Genetics, Public Library of Science, vol. 4(7), pages 1-13, July.
    2. Lei Zhang & Yu-Fang Pei & Jian Li & Christopher J Papasian & Hong-Wen Deng, 2009. "Univariate/Multivariate Genome-Wide Association Scans Using Data from Families and Unrelated Samples," PLOS ONE, Public Library of Science, vol. 4(8), pages 1-12, August.
    3. Wang, Zhendong & Xu, Xingzhong, 2021. "Testing high dimensional covariance matrices via posterior Bayes factor," Journal of Multivariate Analysis, Elsevier, vol. 181(C).
    4. Elaine T. Lim & Yingleong Chan & Pepper Dawes & Xiaoge Guo & Serkan Erdin & Derek J. C. Tai & Songlei Liu & Julia M. Reichert & Mannix J. Burns & Ying Kai Chan & Jessica J. Chiang & Katharina Meyer & , 2022. "Orgo-Seq integrates single-cell and bulk transcriptomic data to identify cell type specific-driver genes associated with autism spectrum disorder," Nature Communications, Nature, vol. 13(1), pages 1-14, December.
    5. Dominic Holland & Oleksandr Frei & Rahul Desikan & Chun-Chieh Fan & Alexey A Shadrin & Olav B Smeland & V S Sundar & Paul Thompson & Ole A Andreassen & Anders M Dale, 2020. "Beyond SNP heritability: Polygenicity and discoverability of phenotypes estimated with a univariate Gaussian mixture model," PLOS Genetics, Public Library of Science, vol. 16(5), pages 1-30, May.
    6. Vincent Michaud & Eulalie Lasseaux & David J. Green & Dave T. Gerrard & Claudio Plaisant & Tomas Fitzgerald & Ewan Birney & Benoît Arveiler & Graeme C. Black & Panagiotis I. Sergouniotis, 2022. "The contribution of common regulatory and protein-coding TYR variants to the genetic architecture of albinism," Nature Communications, Nature, vol. 13(1), pages 1-8, December.
    7. Natalie DeForest & Yuqi Wang & Zhiyi Zhu & Jacqueline S. Dron & Ryan Koesterer & Pradeep Natarajan & Jason Flannick & Tiffany Amariuta & Gina M. Peloso & Amit R. Majithia, 2024. "Genome-wide discovery and integrative genomic characterization of insulin resistance loci using serum triglycerides to HDL-cholesterol ratio as a proxy," Nature Communications, Nature, vol. 15(1), pages 1-17, December.
    8. Makoto Aoshima & Kazuyoshi Yata, 2014. "A distance-based, misclassification rate adjusted classifier for multiclass, high-dimensional data," Annals of the Institute of Statistical Mathematics, Springer;The Institute of Statistical Mathematics, vol. 66(5), pages 983-1010, October.
    9. Parsa Akbari & Dragana Vuckovic & Luca Stefanucci & Tao Jiang & Kousik Kundu & Roman Kreuzhuber & Erik L. Bao & Janine H. Collins & Kate Downes & Luigi Grassi & Jose A. Guerrero & Stephen Kaptoge & Ju, 2023. "A genome-wide association study of blood cell morphology identifies cellular proteins implicated in disease aetiology," Nature Communications, Nature, vol. 14(1), pages 1-19, December.
    10. repec:plo:pgen00:0020137 is not listed on IDEAS
    11. Yata, Kazuyoshi & Aoshima, Makoto, 2013. "PCA consistency for the power spiked model in high-dimensional settings," Journal of Multivariate Analysis, Elsevier, vol. 122(C), pages 334-354.
    12. Jung, Sungkyu & Sen, Arusharka & Marron, J.S., 2012. "Boundary behavior in High Dimension, Low Sample Size asymptotics of PCA," Journal of Multivariate Analysis, Elsevier, vol. 109(C), pages 190-203.
    13. Shivam Sharma & Shashwat Deepali Nagar & Priscilla Pemu & Stephan Zuchner & Leonardo Mariño-Ramírez & Robert Meller & I. King Jordan, 2025. "Genetic ancestry and population structure in the All of Us Research Program cohort," Nature Communications, Nature, vol. 16(1), pages 1-10, December.
    14. Forzani, Liliana & Gieco, Antonella & Tolmasky, Carlos, 2017. "Likelihood ratio test for partial sphericity in high and ultra-high dimensions," Journal of Multivariate Analysis, Elsevier, vol. 159(C), pages 18-38.
    15. Gang Zheng & Zhaohai Li & Mitchell H. Gail & Joseph L. Gastwirth, 2010. "Impact of Population Substructure on Trend Tests for Genetic Case–Control Association Studies," Biometrics, The International Biometric Society, vol. 66(1), pages 196-204, March.
    16. Marie-Claude Babron & Marie de Tayrac & Douglas N Rutledge & Eleftheria Zeggini & Emmanuelle Génin, 2012. "Rare and Low Frequency Variant Stratification in the UK Population: Description and Impact on Association Tests," PLOS ONE, Public Library of Science, vol. 7(10), pages 1-9, October.
    17. Spilimbergo, Antonio & Giuliano, Paola & Tonon, Giovanni, 2006. "Genetic, Cultural and Geographical Distances," CEPR Discussion Papers 5807, C.E.P.R. Discussion Papers.
    18. repec:plo:pone00:0029848 is not listed on IDEAS
    19. Shen, Keren & Yao, Jianfeng & Li, Wai Keung, 2019. "On a spiked model for large volatility matrix estimation from noisy high-frequency data," Computational Statistics & Data Analysis, Elsevier, vol. 131(C), pages 207-221.
    20. Sandosh Padmanabhan & Olle Melander & Toby Johnson & Anna Maria Di Blasio & Wai K Lee & Davide Gentilini & Claire E Hastie & Cristina Menni & Maria Cristina Monti & Christian Delles & Stewart Laing & , 2010. "Genome-Wide Association Study of Blood Pressure Extremes Identifies Variant near UMOD Associated with Hypertension," PLOS Genetics, Public Library of Science, vol. 6(10), pages 1-11, October.
    21. repec:plo:pgen00:1000078 is not listed on IDEAS
    22. Anna Bykhovskaya & Vadim Gorin & Sasha Sodin, 2025. "How weak are weak factors? Uniform inference for signal strength in signal plus noise models," Papers 2507.18554, arXiv.org, revised Feb 2026.
    23. Maïda, M. & Najim, J. & Péché, S., 2007. "Large deviations for weighted empirical mean with outliers," Stochastic Processes and their Applications, Elsevier, vol. 117(10), pages 1373-1403, October.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pgen00:0020190. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosgenetics (email available below). General contact details of provider: https://journals.plos.org/plosgenetics/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.