Indexing Highly Repetitive String Collections, Part I-Reference-Cited by-同舟云学术

Indexing Highly Repetitive String Collections, Part I

Published:2021-04 Issue:2 Volume:54 Page:1-31
ISSN:0360-0300
Container-title:ACM Computing Surveys
language:en
Short-container-title:ACM Comput. Surv.

Author:

Navarro Gonzalo¹

Affiliation:

1. University of Chile, Santiago, Chile

Abstract

Two decades ago, a breakthrough in indexing string collections made it possible to represent them within their compressed space while at the same time offering indexed search functionalities. As this new technology permeated through applications like bioinformatics, the string collections experienced a growth that outperforms Moore’s Law and challenges our ability to handle them even in compressed form. It turns out, fortunately, that many of these rapidly growing string collections are highly repetitive, so that their information content is orders of magnitude lower than their plain size. The statistical compression methods used for classical collections, however, are blind to this repetitiveness, and therefore a new set of techniques has been developed to properly exploit it. The resulting indexes form a new generation of data structures able to handle the huge repetitive string collections that we are facing. In this survey, formed by two parts, we cover the algorithmic developments that have led to these data structures. In this first part, we describe the distinct compression paradigms that have been used to exploit repetitiveness, and the algorithmic techniques that provide direct access to the compressed strings. In the quest for an ideal measure of repetitiveness, we uncover a fascinating web of relations between those measures, as well as the limits up to which the data can be recovered, and up to which direct access to the compressed data can be provided. This is the basic aspect of indexability, which is covered in the second part of this survey.

Funder

Fondecyt

ANID Basal Funds FB0001, Millennium Science Initiative Program

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science,Theoretical Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/3434399

Reference102 articles.

1. Block trees

Cited by 27 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. New string attractor-based complexities for infinite words;Journal of Combinatorial Theory, Series A;2024-11

2. Linear-size suffix tries and linear-size CDAWGs simplified and improved;Acta Informatica;2024-08-23

3. r-indexing the eBWT;Information and Computation;2024-06

4. A survey of BWT variants for string collections;Bioinformatics;2024-05-24

5. Sketching and Streaming for Dictionary Compression;2024 Data Compression Conference (DCC);2024-03-19