Study on reference-based FASTQ genome sequences compression
At a glance
- Citations
- 0
- References
- 10
- Comments
- 0
Abstract
As the cost of genome sequencing decreases, the large amount of genomic data generated brings the storage problem of this massive data. We still have a lot of work to do in the field of specialized data compression of FASTQ files. This paper aims to explore a reference-based lossless compression algorithm for genome sequences in FASTQ format. We propose a compression scheme based on longest matching by using FMD-index to support exact match searching. At the same time, the reverse complementary sequence is used and the insertion, deletion and replacement operations are described effectively to further improve the compression ratio. In comparison with the experimental results of five compressors on seven sets of genome data, the proposed algorithm significantly improves the FASTQ file compression ratios, and is competitive in running time.
Publication details
- DOI
- 10.1145/3523286.3524511
- OpenAlex
- W4281786521
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.