Optimum Search Schemes for Approximate String Matching Using Bidirectional FM-Index-Reference-Cited by-同舟云学术

Optimum Search Schemes for Approximate String Matching Using Bidirectional FM-Index

Published:2018-04-13 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Kianfar Kiavash^ORCID,Pockrandt Christopher,Torkamandi Bahman^ORCID,Luo Haochen^ORCID,Reinert Knut^ORCID

Abstract

AbstractFinding approximate occurrences of a pattern in a text using a full-text index is a central problem in bioinformatics and has been extensively researched. Bidirectional indices have opened new possibilities in this regard allowing the search to start from anywhere within the pattern and extend in both directions. In particular, use of search schemes (partitioning the pattern and searching the pieces in certain orders with given bounds on errors) can yield significant speed-ups. However, finding optimal search schemes is a difficult combinatorial optimization problem.Here for the first time, we propose a mixed integer program (MIP) capable to solve this optimization problem for Hamming distance with given number of pieces. Our experiments show that the optimal search schemes found by our MIP significantly improve the performance of search in bidirectional FM-index upon previous ad-hoc solutions. For example, approximate matching of 101-bp Illumina reads (with two errors) becomes 35 times faster than standard backtracking. Moreover, despite being performed purely in the index, the running time of search using our optimal schemes (for up to two errors) is comparable to the best state-of-the-art aligners, which benefit from combining search in index with in-text verification using dynamic programming. As a result, we anticipate a full-fledged aligner that employs an intelligent combination of search in the bidirectional FM-index using our optimal search schemes and in-text verification using dynamic programming that will outperform today’s best aligners. The development of such an aligner, called FAMOUS (Fast Approximate string Matching using OptimUm search Schemes), is ongoing as our future work.

Publisher

Cold Spring Harbor Laboratory

Reference20 articles.

1. Replacing suffix trees with enhanced suffix arrays

2. Burrows, M. , Wheeler, D.J. : A block-sorting lossless data compression algorithm. Technical Report 124, Digital SRC Research Report (1994)

3. Ferragina, P. , Manzini, G. : Opportunistic data structures with applications. In: FOCS ’00. (2000) 390–398

4. IBM-ILOG: Cplex 12.7.1, https://www.ibm.com/support/knowledgecenter/en/sssa5p_12.7.1/ilog.odms.studio.help/optimization_studio/topics/cos_home.html (Accessed on Nov. 2, 2017).

5. Karkkainen, J. , Na, J.C. : Faster filters for approximate string matching. In: ALENEX ’07. (2007) 84–90

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Porechop_ABI: discovering unknown adapters in ONT sequencing reads for downstream trimming;2022-07-07

2. GenMap: ultra-fast computation of genome mappability;Bioinformatics;2020-04-04

3. GenMap: Fast and Exact Computation of Genome Mappability;2019-04-26