IDEAS home Printed from https://ideas.repec.org/a/nat/natcom/v13y2022i1d10.1038_s41467-022-31724-3.html
   My bibliography  Save this article

Pan-African genome demonstrates how population-specific genome graphs improve high-throughput sequencing data analysis

Author

Listed:
  • H. Serhat Tetikol

    (Seven Bridges Genomics)

  • Deniz Turgut

    (Seven Bridges Genomics)

  • Kubra Narci

    (Seven Bridges Genomics)

  • Gungor Budak

    (Seven Bridges Genomics)

  • Ozem Kalay

    (Seven Bridges Genomics)

  • Elif Arslan

    (Seven Bridges Genomics)

  • Sinem Demirkaya-Budak

    (Seven Bridges Genomics)

  • Alexey Dolgoborodov

    (Seven Bridges Genomics)

  • Duygu Kabakci-Zorlu

    (Seven Bridges Genomics)

  • Vladimir Semenyuk

    (Seven Bridges Genomics)

  • Amit Jain

    (Seven Bridges Genomics)

  • Brandi N. Davis-Dusenbery

    (Seven Bridges Genomics)

Abstract

Graph-based genome reference representations have seen significant development, motivated by the inadequacy of the current human genome reference to represent the diverse genetic information from different human populations and its inability to maintain the same level of accuracy for non-European ancestries. While there have been many efforts to develop computationally efficient graph-based toolkits for NGS read alignment and variant calling, methods to curate genomic variants and subsequently construct genome graphs remain an understudied problem that inevitably determines the effectiveness of the overall bioinformatics pipeline. In this study, we discuss obstacles encountered during graph construction and propose methods for sample selection based on population diversity, graph augmentation with structural variants and resolution of graph reference ambiguity caused by information overload. Moreover, we present the case for iteratively augmenting tailored genome graphs for targeted populations and demonstrate this approach on the whole-genome samples of African ancestry. Our results show that population-specific graphs, as more representative alternatives to linear or generic graph references, can achieve significantly lower read mapping errors and enhanced variant calling sensitivity, in addition to providing the improvements of joint variant calling without the need of computationally intensive post-processing steps.

Suggested Citation

  • H. Serhat Tetikol & Deniz Turgut & Kubra Narci & Gungor Budak & Ozem Kalay & Elif Arslan & Sinem Demirkaya-Budak & Alexey Dolgoborodov & Duygu Kabakci-Zorlu & Vladimir Semenyuk & Amit Jain & Brandi N., 2022. "Pan-African genome demonstrates how population-specific genome graphs improve high-throughput sequencing data analysis," Nature Communications, Nature, vol. 13(1), pages 1-11, December.
  • Handle: RePEc:nat:natcom:v:13:y:2022:i:1:d:10.1038_s41467-022-31724-3
    DOI: 10.1038/s41467-022-31724-3
    as

    Download full text from publisher

    File URL: https://www.nature.com/articles/s41467-022-31724-3
    File Function: Abstract
    Download Restriction: no

    File URL: https://libkey.io/10.1038/s41467-022-31724-3?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. Lasse Maretty & Jacob Malte Jensen & Bent Petersen & Jonas Andreas Sibbesen & Siyang Liu & Palle Villesen & Laurits Skov & Kirstine Belling & Christian Theil Have & Jose M. G. Izarzugaza & Marie Grosj, 2017. "Sequencing and de novo assembly of 150 genomes from Denmark as a population reference," Nature, Nature, vol. 548(7665), pages 87-91, August.
    2. Clare Bycroft & Colin Freeman & Desislava Petkova & Gavin Band & Lloyd T. Elliott & Kevin Sharp & Allan Motyer & Damjan Vukcevic & Olivier Delaneau & Jared O’Connell & Adrian Cortes & Samantha Welsh &, 2018. "The UK Biobank resource with deep phenotyping and genomic data," Nature, Nature, vol. 562(7726), pages 203-209, October.
    3. Hannes P. Eggertsson & Snaedis Kristmundsdottir & Doruk Beyter & Hakon Jonsson & Astros Skuladottir & Marteinn T. Hardarson & Daniel F. Gudbjartsson & Kari Stefansson & Bjarni V. Halldorsson & Pall Me, 2019. "GraphTyper2 enables population-scale genotyping of structural variation using pangenome graphs," Nature Communications, Nature, vol. 10(1), pages 1-8, December.
    4. Konrad J. Karczewski & Laurent C. Francioli & Grace Tiao & Beryl B. Cummings & Jessica Alföldi & Qingbo Wang & Ryan L. Collins & Kristen M. Laricchia & Andrea Ganna & Daniel P. Birnbaum & Laura D. Gau, 2020. "The mutational constraint spectrum quantified from variation in 141,456 humans," Nature, Nature, vol. 581(7809), pages 434-443, May.
    5. Michael P. Snyder & Thomas R. Gingeras & Jill E. Moore & Zhiping Weng & Mark B. Gerstein & Bing Ren & Ross C. Hardison & John A. Stamatoyannopoulos & Brenton R. Graveley & Elise A. Feingold & Michael , 2020. "Perspectives on ENCODE," Nature, Nature, vol. 583(7818), pages 693-698, July.
    6. L. Duncan & H. Shen & B. Gelaye & J. Meijsen & K. Ressler & M. Feldman & R. Peterson & B. Domingue, 2019. "Analysis of polygenic risk score usage and performance in diverse human populations," Nature Communications, Nature, vol. 10(1), pages 1-9, December.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Vincent Michaud & Eulalie Lasseaux & David J. Green & Dave T. Gerrard & Claudio Plaisant & Tomas Fitzgerald & Ewan Birney & Benoît Arveiler & Graeme C. Black & Panagiotis I. Sergouniotis, 2022. "The contribution of common regulatory and protein-coding TYR variants to the genetic architecture of albinism," Nature Communications, Nature, vol. 13(1), pages 1-8, December.
    2. Erik Schoenmakers & Federica Marelli & Helle F. Jørgensen & W. Edward Visser & Carla Moran & Stefan Groeneweg & Carolina Avalos & Sean J. Jurgens & Nichola Figg & Alison Finigan & Neha Wali & Maura Ag, 2023. "Selenoprotein deficiency disorder predisposes to aortic aneurysm formation," Nature Communications, Nature, vol. 14(1), pages 1-14, December.
    3. Xiaoyi Raymond Gao & Marion Chiariglione & Alexander J. Arch, 2022. "Whole-exome sequencing study identifies rare variants and genes associated with intraocular pressure and glaucoma," Nature Communications, Nature, vol. 13(1), pages 1-10, December.
    4. Nicole Deflaux & Margaret Sunitha Selvaraj & Henry Robert Condon & Kelsey Mayo & Sara Haidermota & Melissa A. Basford & Chris Lunt & Anthony A. Philippakis & Dan M. Roden & Joshua C. Denny & Anjene Mu, 2023. "Demonstrating paths for unlocking the value of cloud genomics through cross cohort analysis," Nature Communications, Nature, vol. 14(1), pages 1-10, December.
    5. Ruoyu Tian & Tian Ge & Hyeokmoon Kweon & Daniel B. Rocha & Max Lam & Jimmy Z. Liu & Kritika Singh & Daniel F. Levey & Joel Gelernter & Murray B. Stein & Ellen A. Tsai & Hailiang Huang & Christopher F., 2024. "Whole-exome sequencing in UK Biobank reveals rare genetic architecture for depression," Nature Communications, Nature, vol. 15(1), pages 1-12, December.
    6. Nazia Pathan & Wei Q. Deng & Matteo Di Scipio & Mohammad Khan & Shihong Mao & Robert W. Morton & Ricky Lali & Marie Pigeyre & Michael R. Chong & Guillaume Paré, 2024. "A method to estimate the contribution of rare coding variants to complex trait heritability," Nature Communications, Nature, vol. 15(1), pages 1-16, December.
    7. Magdalena Zimoń & Yunfeng Huang & Anthi Trasta & Aliaksandr Halavatyi & Jimmy Z. Liu & Chia-Yen Chen & Peter Blattmann & Bernd Klaus & Christopher D. Whelan & David Sexton & Sally John & Wolfgang Hube, 2021. "Pairwise effects between lipid GWAS genes modulate lipid plasma levels and cellular uptake," Nature Communications, Nature, vol. 12(1), pages 1-16, December.
    8. Aimee M. Deaton & Aditi Dubey & Lucas D. Ward & Peter Dornbos & Jason Flannick & Elaine Yee & Simina Ticau & Leila Noetzli & Margaret M. Parker & Rachel A. Hoffing & Carissa Willis & Mollie E. Plekan , 2022. "Rare loss of function variants in the hepatokine gene INHBE protect from abdominal obesity," Nature Communications, Nature, vol. 13(1), pages 1-12, December.
    9. Ricky Lali & Michael Chong & Arghavan Omidi & Pedrum Mohammadi-Shemirani & Ann Le & Edward Cui & Guillaume Paré, 2021. "Calibrated rare variant genetic risk scores for complex disease prediction using large exome sequence repositories," Nature Communications, Nature, vol. 12(1), pages 1-15, December.
    10. Carla Márquez-Luna & Steven Gazal & Po-Ru Loh & Samuel S. Kim & Nicholas Furlotte & Adam Auton & Alkes L. Price, 2021. "Incorporating functional priors improves polygenic prediction accuracy in UK Biobank and 23andMe data sets," Nature Communications, Nature, vol. 12(1), pages 1-11, December.
    11. Young Jin Kim & Sanghoon Moon & Mi Yeong Hwang & Sohee Han & Hye-Mi Jang & Jinhwa Kong & Dong Mun Shin & Kyungheon Yoon & Sung Min Kim & Jong-Eun Lee & Anubha Mahajan & Hyun-Young Park & Mark I. McCar, 2022. "The contribution of common and rare genetic variants to variation in metabolic traits in 288,137 East Asians," Nature Communications, Nature, vol. 13(1), pages 1-13, December.
    12. Joel T. Rämö & Tuomo Kiiskinen & Richard Seist & Kristi Krebs & Masahiro Kanai & Juha Karjalainen & Mitja Kurki & Eija Hämäläinen & Paavo Häppölä & Aki S. Havulinna & Heidi Hautakangas & Reedik Mägi &, 2023. "Genome-wide screen of otosclerosis in population biobanks: 27 loci and shared associations with skeletal structure," Nature Communications, Nature, vol. 14(1), pages 1-14, December.
    13. Jeffrey D. Wall & J. Fah Sathirapongsasuti & Ravi Gupta & Asif Rasheed & Radha Venkatesan & Saurabh Belsare & Ramesh Menon & Sameer Phalke & Anuradha Mittal & John Fang & Deepak Tanneeru & Manjari Des, 2023. "South Asian medical cohorts reveal strong founder effects and high rates of homozygosity," Nature Communications, Nature, vol. 14(1), pages 1-11, December.
    14. Andrew D. Grotzinger & Travis T. Mallard & Zhaowen Liu & Jakob Seidlitz & Tian Ge & Jordan W. Smoller, 2023. "Multivariate genomic architecture of cortical thickness and surface area at multiple levels of analysis," Nature Communications, Nature, vol. 14(1), pages 1-13, December.
    15. Alesha A. Hatton & Fei-Fei Cheng & Tian Lin & Ren-Juan Shen & Jie Chen & Zhili Zheng & Jia Qu & Fan Lyu & Sarah E. Harris & Simon R. Cox & Zi-Bing Jin & Nicholas G. Martin & Dongsheng Fan & Grant W. M, 2024. "Genetic control of DNA methylation is largely shared across European and East Asian populations," Nature Communications, Nature, vol. 15(1), pages 1-12, December.
    16. Matthias Wuttke & Eva König & Maria-Alexandra Katsara & Holger Kirsten & Saeed Khomeijani Farahani & Alexander Teumer & Yong Li & Martin Lang & Burulca Göcmen & Cristian Pattaro & Dorothee Günzel & An, 2023. "Imputation-powered whole-exome analysis identifies genes associated with kidney function and disease in the UK Biobank," Nature Communications, Nature, vol. 14(1), pages 1-16, December.
    17. Jiacheng Miao & Hanmin Guo & Gefei Song & Zijie Zhao & Lin Hou & Qiongshi Lu, 2023. "Quantifying portable genetic effects and improving cross-ancestry genetic prediction with GWAS summary statistics," Nature Communications, Nature, vol. 14(1), pages 1-13, December.
    18. David R. Blair & Thomas J. Hoffmann & Joseph T. Shieh, 2022. "Common genetic variation associated with Mendelian disease severity revealed through cryptic phenotype analysis," Nature Communications, Nature, vol. 13(1), pages 1-15, December.
    19. Ananyo Choudhury & Jean-Tristan Brandenburg & Tinashe Chikowore & Dhriti Sengupta & Palwende Romuald Boua & Nigel J. Crowther & Godfred Agongo & Gershim Asiki & F. Xavier Gómez-Olivé & Isaac Kisiangan, 2022. "Meta-analysis of sub-Saharan African studies provides insights into genetic architecture of lipid traits," Nature Communications, Nature, vol. 13(1), pages 1-13, December.
    20. Derek W. Brown & Liam D. Cato & Yajie Zhao & Satish K. Nandakumar & Erik L. Bao & Eugene J. Gardner & Aubrey K. Hubbard & Alexander DePaulis & Thomas Rehling & Lei Song & Kai Yu & Stephen J. Chanock &, 2023. "Shared and distinct genetic etiologies for different types of clonal hematopoiesis," Nature Communications, Nature, vol. 14(1), pages 1-13, December.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:nat:natcom:v:13:y:2022:i:1:d:10.1038_s41467-022-31724-3. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Sonal Shukla or Springer Nature Abstracting and Indexing (email available below). General contact details of provider: http://www.nature.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.