Title
QualComp: a new lossy compressor for quality scores based on rate distortion theory.
Abstract
Next Generation Sequencing technologies have revolutionized many fields in biology by reducing the time and cost required for sequencing. As a result, large amounts of sequencing data are being generated. A typical sequencing data file may occupy tens or even hundreds of gigabytes of disk space, prohibitively large for many users. This data consists of both the nucleotide sequences and per-base quality scores that indicate the level of confidence in the readout of these sequences. Quality scores account for about half of the required disk space in the commonly used FASTQ format (before compression), and therefore the compression of the quality scores can significantly reduce storage requirements and speed up analysis and transmission of sequencing data.In this paper, we present a new scheme for the lossy compression of the quality scores, to address the problem of storage. Our framework allows the user to specify the rate (bits per quality score) prior to compression, independent of the data to be compressed. Our algorithm can work at any rate, unlike other lossy compression algorithms. We envisage our algorithm as being part of a more general compression scheme that works with the entire FASTQ file. Numerical experiments show that we can achieve a better mean squared error (MSE) for small rates (bits per quality score) than other lossy compression schemes. For the organism PhiX, whose assembled genome is known and assumed to be correct, we show that it is possible to achieve a significant reduction in size with little compromise in performance on downstream applications (e.g., alignment).QualComp is an open source software package, written in C and freely available for download at https://sourceforge.net/projects/qualcomp.
Year
DOI
Venue
2013
10.1186/1471-2105-14-187
BMC Bioinformatics
Keywords
Field
DocType
data compression,genome,algorithms,bioinformatics,microarrays,genomics
Lossy compression,Computer science,FASTQ format,Genomics,Software,Bioinformatics,Data compression,Genetics,Data file,Rate–distortion theory,Speedup
Journal
Volume
Issue
ISSN
14
1
1471-2105
Citations 
PageRank 
References 
26
1.38
24
Authors
6
Name
Order
Citations
PageRank
Idoia Ochoa17013.10
Himanshu Asnani211715.39
Dinesh Bharadia382247.06
Mainak Chowdhury48912.74
Tsachy Weissman51192119.50
G Yona665145.52