conference-paper

A Lossless Compression Pipeline for Petabyte-Scale Whole Genome Sequencing Data

Research footprint

At a glance

Citations
0
References
13
Comments
0
Paper overview

Abstract

Whole genome sequencing (WGS) technologies have enabled high-throughput cost-effective genome sequencing at the population scale. A single WGS instrument can sequence millions of DNA molecules simultaneously, leading to the generation of massive datasets. GenomeIndia is an ongoing national project aimed at sequencing the genomes of 10,000 Indian individuals. The GenomeIndia sequencing centers are completing the generation of petabyte-scale genomic data. This has raised an urgent need for scalable lossless compression software to facilitate cost-effective storage and exchange of data. By default, each WGS file produced in the GenomeIndia project is stored in the standard unmapped BAM (uBAM) format. A uBAM file saves the DNA sequences as well as metadata associated with the sequencing experiment. We have developed an open-source software pipeline that enables parallel lossless compression and decompression of uBAM files. It produces compressed output that is approximately 5 x smaller than the input uBAM files. We carefully engineered the pipeline by integrating different bioinformatics tools such as SPRING, Picard, SAMtools, and PySAM. We evaluated the parallel efficiency of our approach using thorough performance profiling and strong-scaling experiments.

Record transparency

Publication details

DOI
10.1109/hipc58850.2023.00056
OpenAlex
W4393973379
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.