article Open access

gsufsort: constructing suffix arrays, LCP arrays and BWTs for string collections

  • Algorithms for Molecular Biology
  • BioMed Central
Research footprint

At a glance

Citations
19
References
21
Comments
0
Paper overview

Abstract

BACKGROUND: The construction of a suffix array for a collection of strings is a fundamental task in Bioinformatics and in many other applications that process strings. Related data structures, as the Longest Common Prefix array, the Burrows-Wheeler transform, and the document array, are often needed to accompany the suffix array to efficiently solve a wide variety of problems. While several algorithms have been proposed to construct the suffix array for a single string, less emphasis has been put on algorithms to construct suffix arrays for string collections. RESULT: ) time. Our tool is written in ANSI/C and is based on the algorithm gSACA-K (Louza et al. in Theor Comput Sci 678:22-39, 2017), the fastest algorithm to construct suffix arrays for string collections. The tool supports large fasta, fastq and text files with multiple strings as input. Experiments have shown very good performance on different types of strings. CONCLUSIONS: gsufsort is a fast, portable, and lightweight tool for constructing the suffix array and additional data structures for string collections.

Record transparency

Publication details

DOI
10.1186/s13015-020-00177-y
OpenAlex
W3089043377
Document type
article
Language
EN
Source
Algorithms for Molecular Biology
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.