Fractional Hitting Sets for Efficient and Lightweight Genomic Data Sketching-Reference-Cited by-同舟云学术

Fractional Hitting Sets for Efficient and Lightweight Genomic Data Sketching

Published:2023-06-24 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Rouzé Timothé^ORCID,Martayan Igor^ORCID,Marchet Camille^ORCID,Limasset Antoine^ORCID

Abstract

AbstractThe exponential increase in publicly available sequencing data and genomic resources necessitates the development of highly efficient methods for data processing and analysis. Locality-sensitive hashing techniques have successfully transformed large datasets into smaller, more manageable sketches while maintaining comparability using metrics such as Jaccard and containment indices. However, fixed-size sketches encounter difficulties when applied to divergent datasets.Scalable sketching methods, such as Sourmash, provide valuable solutions but still lack resourceefficient, tailored indexing. Our objective is to create lighter sketches with comparable results while enhancing efficiency. We introduce the concept of Fractional Hitting Sets, a generalization of Universal Hitting Sets, which uniformly cover a specified fraction of thek-mer space. In theory and practice, we demonstrate the feasibility of achieving such coverage with simple but highly efficient schemes.By encoding the coveredk-mers as super-k-mers, we provide a space-efficient exact representation that also enables optimized comparisons. Our novel tool, SuperSampler, implements this scheme, and experimental results with real bacterial collections closely match our theoretical findings.In comparison to Sourmash, SuperSampler achieves similar outcomes while utilizing an order of magnitude less space and memory and operating several times faster. This highlights the potential of our approach in addressing the challenges presented by the ever-expanding landscape of genomic data.SuperSampler is an open-source software and can be accessed atgithub.com/TimRouze/supersampler. The data required to reproduce the results presented in this manuscript is available atgithub.com/TimRouze/Expe_SPSP.

Publisher

Cold Spring Harbor Laboratory

Reference31 articles.

1. Clément Agret , Bastien Cazaux , and Antoine Limasset . Toward optimal fingerprint indexing for large scale genomics. In 22nd International Workshop on Algorithms in Bioinformatics, 2022.

2. Performance of neural network basecalling tools for Oxford Nanopore sequencing

3. Daniel N Baker and Ben Langmead . Dashing 2: genomic sketching with multiplicities and locality-sensitive hashing. In RECOMB, 2023.

4. Multiple comparative metagenomics using multiset k-mer counting;PeerJ Computer Science,2016

5. Exploring bacterial diversity via a curated and searchable snapshot of archived dna sequences;PLoS biology,2021

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A survey of k-mer methods and applications in bioinformatics;Computational and Structural Biotechnology Journal;2024-12

2. k-nonical space: sketching with reverse complements;2024-01-27