Title
Compressive biological sequence analysis and archival in the era of high-throughput sequencing technologies.
Abstract
High-throughput sequencing technologies produce large collections of data, mainly DNA sequences with additional information, requiring the design of efficient and effective methodologies for both their compression and storage. In this context, we first provide a classification of the main techniques that have been proposed, according to three specific research directions that have emerged from the literature and, for each, we provide an overview of the current techniques. Finally, to make this review useful to researchers and technicians applying the existing software and tools, we include a synopsis of the main characteristics of the described approaches, including details on their implementation and availability. Performance of the various methods is also highlighted, although the state of the art does not lend itself to a consistent and coherent comparison among all the methods presented here.
Year
DOI
Venue
2014
10.1093/bib/bbt088
BRIEFINGS IN BIOINFORMATICS
Keywords
DocType
Volume
data compression of large sequence collections,data compression in bioinformatics,storage and management of HTS data,compressive sequence analysis,analysis of large biological sequence collections,succinct data structures for bioinformatics
Journal
15
Issue
ISSN
Citations 
SP3
1467-5463
5
PageRank 
References 
Authors
0.46
36
3
Name
Order
Citations
PageRank
Raffaele Giancarlo11112107.23
Simona E. Rombo219222.21
F. Utro320416.63