IDEAS home Printed from https://ideas.repec.org/a/plo/pgen00/1007978.html
   My bibliography  Save this article

Fast and flexible linear mixed models for genome-wide genetics

Author

Listed:
  • Daniel E Runcie
  • Lorin Crawford

Abstract

Linear mixed effect models are powerful tools used to account for population structure in genome-wide association studies (GWASs) and estimate the genetic architecture of complex traits. However, fully-specified models are computationally demanding and common simplifications often lead to reduced power or biased inference. We describe Grid-LMM (https://github.com/deruncie/GridLMM), an extendable algorithm for repeatedly fitting complex linear models that account for multiple sources of heterogeneity, such as additive and non-additive genetic variance, spatial heterogeneity, and genotype-environment interactions. Grid-LMM can compute approximate (yet highly accurate) frequentist test statistics or Bayesian posterior summaries at a genome-wide scale in a fraction of the time compared to existing general-purpose methods. We apply Grid-LMM to two types of quantitative genetic analyses. The first is focused on accounting for spatial variability and non-additive genetic variance while scanning for QTL; and the second aims to identify gene expression traits affected by non-additive genetic variation. In both cases, modeling multiple sources of heterogeneity leads to new discoveries.Author summary: The goal of quantitative genetics is to characterize the relationship between genetic variation and variation in quantitative traits such as height, productivity, or disease susceptibility. A statistical method known as the linear mixed effect model has been critical to the development of quantitative genetics. First applied to animal breeding, this model now forms the basis of a wide-range of modern genomic analyses including genome-wide associations, polygenic modeling, and genomic prediction. The same model is also widely used in ecology, evolutionary genetics, social sciences, and many other fields. Mixed models are frequently multi-faceted, which is necessary for accurately modeling data that is generated from complex experimental designs. However, most genomic applications use only the simplest form of linear mixed methods because the computational demands for model fitting can be too great. We develop a flexible approach for fitting linear mixed models to genome scale data that greatly reduces their computational burden and provides flexibility for users to choose the best statistical paradigm for their data analysis. We demonstrate improved accuracy for genetic association tests, increased power to discover causal genetic variants, and the ability to provide accurate summaries of model uncertainty using both simulated and real data examples.

Suggested Citation

  • Daniel E Runcie & Lorin Crawford, 2019. "Fast and flexible linear mixed models for genome-wide genetics," PLOS Genetics, Public Library of Science, vol. 15(2), pages 1-24, February.
  • Handle: RePEc:plo:pgen00:1007978
    DOI: 10.1371/journal.pgen.1007978
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.1007978
    Download Restriction: no

    File URL: https://journals.plos.org/plosgenetics/article/file?id=10.1371/journal.pgen.1007978&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pgen.1007978?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. B. Devlin & Kathryn Roeder, 1999. "Genomic Control for Association Studies," Biometrics, The International Biometric Society, vol. 55(4), pages 997-1004, December.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Odín Morón-García & Gina A Garzón-Martínez & M J Pilar Martínez-Martín & Jason Brook & Fiona M K Corke & John H Doonan & Anyela V Camargo Rodríguez, 2022. "Genetic architecture of variation in Arabidopsis thaliana rosettes," PLOS ONE, Public Library of Science, vol. 17(2), pages 1-22, February.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Lei Zhang & Yu-Fang Pei & Jian Li & Christopher J Papasian & Hong-Wen Deng, 2009. "Univariate/Multivariate Genome-Wide Association Scans Using Data from Families and Unrelated Samples," PLOS ONE, Public Library of Science, vol. 4(8), pages 1-12, August.
    2. Dominic Holland & Oleksandr Frei & Rahul Desikan & Chun-Chieh Fan & Alexey A Shadrin & Olav B Smeland & V S Sundar & Paul Thompson & Ole A Andreassen & Anders M Dale, 2020. "Beyond SNP heritability: Polygenicity and discoverability of phenotypes estimated with a univariate Gaussian mixture model," PLOS Genetics, Public Library of Science, vol. 16(5), pages 1-30, May.
    3. Vincent Michaud & Eulalie Lasseaux & David J. Green & Dave T. Gerrard & Claudio Plaisant & Tomas Fitzgerald & Ewan Birney & Benoît Arveiler & Graeme C. Black & Panagiotis I. Sergouniotis, 2022. "The contribution of common regulatory and protein-coding TYR variants to the genetic architecture of albinism," Nature Communications, Nature, vol. 13(1), pages 1-8, December.
    4. Parsa Akbari & Dragana Vuckovic & Luca Stefanucci & Tao Jiang & Kousik Kundu & Roman Kreuzhuber & Erik L. Bao & Janine H. Collins & Kate Downes & Luigi Grassi & Jose A. Guerrero & Stephen Kaptoge & Ju, 2023. "A genome-wide association study of blood cell morphology identifies cellular proteins implicated in disease aetiology," Nature Communications, Nature, vol. 14(1), pages 1-19, December.
    5. Gang Zheng & Zhaohai Li & Mitchell H. Gail & Joseph L. Gastwirth, 2010. "Impact of Population Substructure on Trend Tests for Genetic Case–Control Association Studies," Biometrics, The International Biometric Society, vol. 66(1), pages 196-204, March.
    6. Sandosh Padmanabhan & Olle Melander & Toby Johnson & Anna Maria Di Blasio & Wai K Lee & Davide Gentilini & Claire E Hastie & Cristina Menni & Maria Cristina Monti & Christian Delles & Stewart Laing & , 2010. "Genome-Wide Association Study of Blood Pressure Extremes Identifies Variant near UMOD Associated with Hypertension," PLOS Genetics, Public Library of Science, vol. 6(10), pages 1-11, October.
    7. Jakris Eu-ahsunthornwattana & E Nancy Miller & Michaela Fakiola & Wellcome Trust Case Control Consortium 2 & Selma M B Jeronimo & Jenefer M Blackwell & Heather J Cordell, 2014. "Comparison of Methods to Account for Relatedness in Genome-Wide Association Studies with Family-Based Data," PLOS Genetics, Public Library of Science, vol. 10(7), pages 1-20, July.
    8. Jianzhong Ma & Christopher I Amos, 2010. "Theoretical Formulation of Principal Components Analysis to Detect and Correct for Population Stratification," PLOS ONE, Public Library of Science, vol. 5(9), pages 1-14, September.
    9. Claire L Simpson & Robert Wojciechowski & Konrad Oexle & Federico Murgia & Laura Portas & Xiaohui Li & Virginie J M Verhoeven & Veronique Vitart & Maria Schache & S Mohsen Hosseini & Pirro G Hysi & Le, 2014. "Genome-Wide Meta-Analysis of Myopia and Hyperopia Provides Evidence for Replication of 11 Loci," PLOS ONE, Public Library of Science, vol. 9(9), pages 1-19, September.
    10. Matthieu Bouaziz & Christophe Ambroise & Mickael Guedj, 2011. "Accounting for Population Stratification in Practice: A Comparison of the Main Strategies Dedicated to Genome-Wide Association Studies," PLOS ONE, Public Library of Science, vol. 6(12), pages 1-13, December.
    11. Aditi Shendre & Howard W Wiener & Marguerite R Irvin & Bradley E Aouizerat & Edgar T Overton & Jason Lazar & Chenglong Liu & Howard N Hodis & Nita A Limdi & Kathleen M Weber & Stephen J Gange & Degui , 2017. "Genome-wide admixture and association study of subclinical atherosclerosis in the Women’s Interagency HIV Study (WIHS)," PLOS ONE, Public Library of Science, vol. 12(12), pages 1-23, December.
    12. Li Shaoyu & Lu Qing & Fu Wenjiang & Romero Roberto & Cui Yuehua, 2009. "A Regularized Regression Approach for Dissecting Genetic Conflicts that Increase Disease Risk in Pregnancy," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 8(1), pages 1-28, October.
    13. Warrington Nicole M. & Tilling Kate & Howe Laura D. & Paternoster Lavinia & Pennell Craig E. & Wu Yan Yan & Briollais Laurent, 2014. "Robustness of the linear mixed effects model to error distribution assumptions and the consequences for genome-wide association studies," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 13(5), pages 1-21, October.
    14. Wang, Linglu & Li, Qizhai & Li, Zhaohai & Zheng, Gang, 2011. "Bayes factors in the presence of population stratification," Statistics & Probability Letters, Elsevier, vol. 81(7), pages 836-841, July.
    15. Boitard Simon & Mangin Brigitte & Azaïs Jean-Marc, 2010. "Asymptotic Distribution of the "Orthogonal" Quantitative Transmission Disequilibrium Test in a Structured Population: Exact Formula," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 9(1), pages 1-25, January.
    16. Ilja M Nolte & Chris Wallace & Stephen J Newhouse & Daryl Waggott & Jingyuan Fu & Nicole Soranzo & Rhian Gwilliam & Panos Deloukas & Irina Savelieva & Dongling Zheng & Chrysoula Dalageorgou & Martin F, 2009. "Common Genetic Variation Near the Phospholamban Gene Is Associated with Cardiac Repolarisation: Meta-Analysis of Three Genome-Wide Association Studies," PLOS ONE, Public Library of Science, vol. 4(7), pages 1-10, July.
    17. Nick Patterson & Alkes L Price & David Reich, 2006. "Population Structure and Eigenanalysis," PLOS Genetics, Public Library of Science, vol. 2(12), pages 1-20, December.
    18. Ferguson John P. & Palejev Dean, 2014. "P-value calibration for multiple testing problems in genomics," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 13(6), pages 1-15, December.
    19. Tiago C. Silva & Juan I. Young & Lanyu Zhang & Lissette Gomez & Michael A. Schmidt & Achintya Varma & X. Steven Chen & Eden R. Martin & Lily Wang, 2022. "Cross-tissue analysis of blood and brain epigenome-wide association studies in Alzheimer’s disease," Nature Communications, Nature, vol. 13(1), pages 1-16, December.
    20. Takeshi Nishiyama & Hirohisa Kishino & Sadao Suzuki & Ryosuke Ando & Hideshi Niimura & Hirokazu Uemura & Mikako Horita & Keizo Ohnaka & Nagato Kuriyama & Haruo Mikami & Naoyuki Takashima & Keitaro Mas, 2012. "Detailed Analysis of Japanese Population Substructure with a Focus on the Southwest Islands of Japan," PLOS ONE, Public Library of Science, vol. 7(4), pages 1-7, April.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pgen00:1007978. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosgenetics (email available below). General contact details of provider: https://journals.plos.org/plosgenetics/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.