On the Optimal Design of Genetic Variant Discovery Studies
The recent emergence of massively parallel sequencing technologies has enabled an increasing number of human genome re-sequencing studies, notable among them being the 1000 Genomes Project. The main aim of these studies is to identify the yet unknown genetic variants in a genomic region, mostly low frequency variants (frequency less than 5%). We propose here a set of statistical tools that address how to optimally design such studies in order to increase the number of genetic variants we expect to discover. Within this framework, the tradeoff between lower coverage for more individuals and higher coverage for fewer individuals can be naturally solved.The methods here are also useful for estimating the number of genetic variants missed in a discovery study performed at low coverage.We show applications to simulated data based on coalescent models and to sequence data from the ENCODE project. In particular, we show the extent to which combining data from multiple populations in a discovery study may increase the number of genetic variants identified relative to studies on single populations.
Volume (Year): 9 (2010)
Issue (Month): 1 (August)
|Contact details of provider:|| Web page: http://www.degruyter.com|
|Order Information:||Web: http://www.degruyter.com/view/j/sagmb|
When requesting a correction, please mention this item's handle: RePEc:bpj:sagmbi:v:9:y:2010:i:1:n:33. See general information about how to correct material in RePEc.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: (Peter Golla)
If references are entirely missing, you can add them using this form.