Abstract
(1) Background: DNA sequence alignment process is an essential step in genome analysis. BWA-MEM has been a prevalent single-node tool in genome alignment because of its high speed and accuracy. The exponentially generated genome data requiring a multi-node solution to handle large volumes of data currently remains a challenge. Spark is a ubiquitous big data platform that has been exploited to assist genome alignment in handling this challenge. Nonetheless, existing works that utilize Spark to optimize BWA-MEM suffer from higher overhead. (2) Methods: In this paper, we presented PipeMEM, a framework to accelerate BWA-MEM with lower overhead with the help of the pipe operation in Spark. We additionally proposed to use a pipeline structure and in-memory-computation to accelerate PipeMEM. (3) Results: Our experiments showed that, on paired-end alignment tasks, our framework had low overhead. In a multi-node environment, our framework, on average, was 2.27× faster compared with BWASpark (an alignment tool in Genome Analysis Toolkit (GATK)), and 2.33× faster compared with SparkBWA. (4) Conclusions: PipeMEM could accelerate BWA-MEM in the Spark environment with high performance and low overhead.
Funder
Guangdong Natural Science Foundation
Subject
Genetics(clinical),Genetics
Cited by
11 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献
1. Bioinformatics characterization of variants of uncertain significance in pediatric sensorineural hearing loss;Frontiers in Pediatrics;2024-02-21
2. Repeats in Genomes;Reference Module in Life Sciences;2024
3. Efficient Variant Calling on Human Genome Sequences Using a GPU-Enabled Commodity Cluster;Proceedings of the 32nd ACM International Conference on Information and Knowledge Management;2023-10-21
4. SparkFlow: Towards High-Performance Data Analytics for Spark-based Genome Analysis;2022 22nd IEEE International Symposium on Cluster, Cloud and Internet Computing (CCGrid);2022-05
5. Analytical Pipelines for the GBS Analysis;Genotyping by Sequencing for Crop Improvement;2022-04