Title
Haplotype assembly in polyploid genomes and identical by descent shared tracts.
Abstract
Motivation: Genome-wide haplotype reconstruction from sequence data, or haplotype assembly, is at the center of major challenges in molecular biology and life sciences. For complex eukaryotic organisms like humans, the genome is vast and the population samples are growing so rapidly that algorithms processing high-throughput sequencing data must scale favorably in terms of both accuracy and computational efficiency. Furthermore, current models and methodologies for haplotype assembly (i) do not consider individuals sharing haplotypes jointly, which reduces the size and accuracy of assembled haplotypes, and (ii) are unable to model genomes having more than two sets of homologous chromosomes (polyploidy). Polyploid organisms are increasingly becoming the target of many research groups interested in the genomics of disease, phylogenetics, botany and evolution but there is an absence of theory and methods for polyploid haplotype reconstruction. Results: In this work, we present a number of results, extensions and generalizations of compass graphs and our HapCompass framework. We prove the theoretical complexity of two haplotype assembly optimizations, thereby motivating the use of heuristics. Furthermore, we present graph theory-based algorithms for the problem of haplotype assembly using our previously developed HapCompass framework for (i) novel implementations of haplotype assembly optimizations (minimum error correction), (ii) assembly of a pair of individuals sharing a haplotype tract identical by descent and (iii) assembly of polyploid genomes. We evaluate our methods on 1000 Genomes Project, Pacific Biosciences and simulated sequence data.
Year
DOI
Venue
2013
10.1093/bioinformatics/btt213
BIOINFORMATICS
Keywords
Field
DocType
haplotypes,algorithms,genomics
Genome,Population,Identity by descent,Computer science,Haplotype,Genomics,1000 Genomes Project,Bioinformatics,Phylogenetics,Haplotype estimation
Journal
Volume
Issue
ISSN
29
13
1367-4803
Citations 
PageRank 
References 
20
1.15
14
Authors
2
Name
Order
Citations
PageRank
Derek Aguiar1675.18
Sorin Istrail21415170.40