A Lossless Compression Pipeline for Petabyte-Scale Whole Genome Sequencing Data
At a glance
- Citations
- 0
- References
- 13
- Comments
- 0
Abstract
Whole genome sequencing (WGS) technologies have enabled high-throughput cost-effective genome sequencing at the population scale. A single WGS instrument can sequence millions of DNA molecules simultaneously, leading to the generation of massive datasets. GenomeIndia is an ongoing national project aimed at sequencing the genomes of 10,000 Indian individuals. The GenomeIndia sequencing centers are completing the generation of petabyte-scale genomic data. This has raised an urgent need for scalable lossless compression software to facilitate cost-effective storage and exchange of data. By default, each WGS file produced in the GenomeIndia project is stored in the standard unmapped BAM (uBAM) format. A uBAM file saves the DNA sequences as well as metadata associated with the sequencing experiment. We have developed an open-source software pipeline that enables parallel lossless compression and decompression of uBAM files. It produces compressed output that is approximately 5 x smaller than the input uBAM files. We carefully engineered the pipeline by integrating different bioinformatics tools such as SPRING, Picard, SAMtools, and PySAM. We evaluated the parallel efficiency of our approach using thorough performance profiling and strong-scaling experiments.
Publication details
- DOI
- 10.1109/hipc58850.2023.00056
- OpenAlex
- W4393973379
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.