Optimizing performance of GATK workflows using Apache Arrow In-Memory data framework-Reference-Cited by-同舟云学术

Optimizing performance of GATK workflows using Apache Arrow In-Memory data framework

Published:2020-11 Issue:S10 Volume:21 Page:
ISSN:1471-2164
Container-title:BMC Genomics
language:en
Short-container-title:BMC Genomics

Author:

Ahmad Tanveer,Ahmed Nauman,Al-Ars Zaid,Hofstee H. Peter

Abstract

Abstract Background Immense improvements in sequencing technologies enable producing large amounts of high throughput and cost effective next-generation sequencing (NGS) data. This data needs to be processed efficiently for further downstream analyses. Computing systems need this large amounts of data closer to the processor (with low latency) for fast and efficient processing. However, existing workflows depend heavily on disk storage and access, to process this data incurs huge disk I/O overheads. Previously, due to the cost, volatility and other physical constraints of DRAM memory, it was not feasible to place large amounts of working data sets in memory. However, recent developments in storage-class memory and non-volatile memory technologies have enabled computing systems to place huge data in memory to process it directly from memory to avoid disk I/O bottlenecks. To exploit the benefits of such memory systems efficiently, proper formatted data placement in memory and its high throughput access is necessary by avoiding (de)-serialization and copy overheads in between processes. For this purpose, we use the newly developed Apache Arrow, a cross-language development framework that provides language-independent columnar in-memory data format for efficient in-memory big data analytics. This allows genomics applications developed in different programming languages to communicate in-memory without having to access disk storage and avoiding (de)-serialization and copy overheads. Implementation We integrate Apache Arrow in-memory based Sequence Alignment/Map (SAM) format and its shared memory objects store library in widely used genomics high throughput data processing applications like BWA-MEM, Picard and GATK to allow in-memory communication between these applications. In addition, this also allows us to exploit the cache locality of tabular data and parallel processing capabilities through shared memory objects. Results Our implementation shows that adopting in-memory SAM representation in genomics high throughput data processing applications results in better system resource utilization, low number of memory accesses due to high cache locality exploitation and parallel scalability due to shared memory objects. Our implementation focuses on the GATK best practices recommended workflows for germline analysis on whole genome sequencing (WGS) and whole exome sequencing (WES) data sets. We compare a number of existing in-memory data placing and sharing techniques like ramDisk and Unix pipes to show how columnar in-memory data representation outperforms both. We achieve a speedup of 4.85x and 4.76x for WGS and WES data, respectively, in overall execution time of variant calling workflows. Similarly, a speedup of 1.45x and 1.27x for these data sets, respectively, is achieved, as compared to the second fastest workflow. In some individual tools, particularly in sorting, duplicates removal and base quality score recalibration the speedup is even more promising. Availability The code and scripts used in our experiments are available in both container and repository form at: https://github.com/abs-tudelft/ArrowSAM.

Publisher

Springer Science and Business Media LLC

Subject

Genetics,Biotechnology

Link

http://link.springer.com/content/pdf/10.1186/s12864-020-07013-y.pdf

Reference52 articles.

1. Xia X. Comparative Genomics; 2013. https://doi.org/10.1007/978-3-642-37146-2.

2. Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. J Mol Biol. 1990; 215(3):403–10. https://doi.org/10.1016/S0022-2836(05)80360-2.

3. J Lipman D, Pearson W. Rapid and sensitive protein similarity searches. Science (New York, N.Y.) 1985; 227:1435–41. https://doi.org/10.1126/science.2983426.

4. Wheeler WC, S. Gladstein D. Malign: A multiple sequence alignment program. J Hered. 1994; 85. https://doi.org/10.1093/oxfordjournals.jhered.a111492.

5. Rice P, Longden I, Bleasby A. Emboss: The european molecular biology open software suite. Trends Genet TIG. 2000; 16:276–7. https://doi.org/10.1016/S0168-9525(00)02024-2.

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Dengue Virus Surveillance in Nepal Yields the First On-Site Whole Genome Sequences of Isolates from the 2022 Outbreak;2024-06-03

2. COSAP: Comparative Sequencing Analysis Platform;BMC Bioinformatics;2024-03-26

3. Communication-Efficient Cluster Scalable Genomics Data Processing Using Apache Arrow Flight;2022 21st International Symposium on Parallel and Distributed Computing (ISPDC);2022-07

4. REIP: A Reconfigurable Environmental Intelligence Platform and Software Framework for Fast Sensor Network Prototyping;Sensors;2022-05-17

5. Data Integration Challenges for Machine Learning in Precision Medicine;Frontiers in Medicine;2022-01-25