Title
Integrating sequencing technologies in personal genomics: optimal low cost reconstruction of structural variants.
Abstract
The goal of human genome re-sequencing is obtaining an accurate assembly of an individual's genome. Recently, there has been great excitement in the development of many technologies for this (e.g. medium and short read sequencing from companies such as 454 and SOLiD, and high-density oligo-arrays from Affymetrix and NimbelGen), with even more expected to appear. The costs and sensitivities of these technologies differ considerably from each other. As an important goal of personal genomics is to reduce the cost of re-sequencing to an affordable point, it is worthwhile to consider optimally integrating technologies. Here, we build a simulation toolbox that will help us optimally combine different technologies for genome re-sequencing, especially in reconstructing large structural variants (SVs). SV reconstruction is considered the most challenging step in human genome re-sequencing. (It is sometimes even harder than de novo assembly of small genomes because of the duplications and repetitive sequences in the human genome.) To this end, we formulate canonical problems that are representative of issues in reconstruction and are of small enough scale to be computationally tractable and simulatable. Using semi-realistic simulations, we show how we can combine different technologies to optimally solve the assembly at low cost. With mapability maps, our simulations efficiently handle the inhomogeneous repeat-containing structure of the human genome and the computational complexity of practical assembly algorithms. They quantitatively show how combining different read lengths is more cost-effective than using one length, how an optimal mixed sequencing strategy for reconstructing large novel SVs usually also gives accurate detection of SNPs/indels, how paired-end reads can improve reconstruction efficiency, and how adding in arrays is more efficient than just sequencing for disentangling some complex SVs. Our strategy should facilitate the sequencing of human genomes at maximum accuracy and low cost.
Year
DOI
Venue
2009
10.1371/journal.pcbi.1000432
PLOS COMPUTATIONAL BIOLOGY
Keywords
Field
DocType
cost effectiveness,genomics,computational complexity,computer simulation,human genome
Genome,Hybrid genome assembly,Biology,Comparative genomics,Genomics,Genome evolution,DNA sequencing,Bioinformatics,Genetics,Personal genomics,Sequence assembly
Journal
Volume
Issue
ISSN
5
7
1553-734X
Citations 
PageRank 
References 
3
0.84
3
Authors
6
Name
Order
Citations
PageRank
Jiang Du1182.62
Robert D. Bjornson241.55
Zhengdong D. Zhang3426.96
Yong Kong4243.27
Michael Snyder513826.15
Mark B. Gerstein623918.21