Improving on hash-based probabilistic sequence classification using multiple spaced seeds and multi-index Bloom filters-Reference-Cited by-同舟云学术

Improving on hash-based probabilistic sequence classification using multiple spaced seeds and multi-index Bloom filters

Published:2018-10-05 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Chu Justin^ORCID,Mohamadi Hamid,Erhan Emre,Tse Jeffery,Chiu Readman,Yeo Sarah,Birol Inanc

Abstract

ABSTRACTAlignment-free classification of sequences against collections of sequences has enabled high-throughput processing of sequencing data in many bioinformatics analysis pipelines. Originally hash-table based, much work has been done to improve and reduce the memory requirement of indexing of k-mer sequences with probabilistic indexing strategies. These efforts have led to lower memory highly efficient indexes, but often lack sensitivity in the face of sequencing errors or polymorphism because they are k-mer based. To address this, we designed a new memory efficient data structure that can tolerate mismatches using multiple spaced seeds, called a multi-index Bloom Filter. Implemented as part of BioBloom Tools, we demonstrate our algorithm in two applications, read binning for targeted assembly and taxonomic read assignment. Our tool shows a higher sensitivity and specificity for read-binning than BWA MEM at an order of magnitude less time. For taxonomic classification, we show higher sensitivity than CLARK-S at an order of magnitude less time while using half the memory.

Publisher

Cold Spring Harbor Laboratory

Reference48 articles.

1. PatternHunter: faster and more sensitive homology search

2. Burkhardt, S. and Kärkkäinen, J. (2002) Annual Symposium on Combinatorial Pattern Matching. Springer, pp. 225–234.

3. Enhanced Regulatory Sequence Prediction Using Gapped k-mer Features

4. A coverage criterion for spaced seeds and its applications to support vector machine string kernels and k-mer distances;Journal of computational biology : a journal of computational molecular cell biology,2014

5. BOND: Basic OligoNucleotide Design;BMC bioinformatics,2013

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Data Structures to Represent a Set of k -long DNA Sequences;ACM Computing Surveys;2021-04

2. To Petabytes and beyond: recent advances in probabilistic and signal processing algorithms and their application to metagenomics;Nucleic Acids Research;2020-04-27