Algorithm-Based Fault Tolerance for Dense Matrix Factorizations, Multiple Failures and Accuracy-Reference-Cited by-同舟云学术

Algorithm-Based Fault Tolerance for Dense Matrix Factorizations, Multiple Failures and Accuracy

Published:2015-02-18 Issue:2 Volume:1 Page:1-28
ISSN:2329-4949
Container-title:ACM Transactions on Parallel Computing
language:en
Short-container-title:ACM Trans. Parallel Comput.

Author:

Bouteiller Aurelien¹,Herault Thomas¹,Bosilca George¹,Du Peng¹,Dongarra Jack¹

Affiliation:

1. Innovative Computing Laboratory, University of Tennessee, Knoxville

Abstract

Dense matrix factorizations, such as LU, Cholesky and QR, are widely used for scientific applications that require solving systems of linear equations, eigenvalues and linear least squares problems. Such computations are normally carried out on supercomputers, whose ever-growing scale induces a fast decline of the Mean Time To Failure (MTTF). This article proposes a new hybrid approach, based on Algorithm-Based Fault Tolerance (ABFT), to help matrix factorizations algorithms survive fail-stop failures. We consider extreme conditions, such as the absence of any reliable node and the possibility of losing both data and checksum from a single failure. We will present a generic solution for protecting the right factor, where the updates are applied, of all above mentioned factorizations. For the left factor, where the panel has been applied, we propose a scalable checkpointing algorithm. This algorithm features high degree of checkpointing parallelism and cooperatively utilizes the checksum storage leftover from the right factor protection. The fault-tolerant algorithms derived from this hybrid solution is applicable to a wide range of dense matrix factorizations, with minor modifications. Theoretical analysis shows that the fault tolerance overhead decreases inversely to the scaling in the number of computing units and the problem size. Experimental results of LU and QR factorization on the Kraken (Cray XT5) supercomputer validate the theoretical evaluation and confirm negligible overhead, with- and without-errors. Applicability to tolerate multiple failures and accuracy after multiple recovery is also considered.

Funder

U.S. Department of Energy

National Science Foundation

Publisher

Association for Computing Machinery (ACM)

Subject

Computational Theory and Mathematics,Computer Science Applications,Hardware and Architecture,Modeling and Simulation,Software

Link

https://dl.acm.org/doi/pdf/10.1145/2686892

Reference23 articles.

1. Algorithm-based fault tolerance applied to high performance computing

2. Fault Tolerance in Petascale/ Exascale Systems: Current Knowledge, Challenges and Research Opportunities

Cited by 15 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Taking the MPI standard and the open MPI library to exascale;The International Journal of High Performance Computing Applications;2024-07-23

2. Rollback-Free Recovery for a High Performance Dense Linear Solver With Reduced Memory Footprint;IEEE Transactions on Parallel and Distributed Systems;2024-07

3. Software approaches for resilience of high performance computing systems: a survey;Frontiers of Computer Science;2022-12-12

4. Implicit Actions and Non-blocking Failure Recovery with MPI;2022 IEEE/ACM 12th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS);2022-11

5. Performance and power modeling and prediction using MuMMI and 10 machine learning methods;Concurrency and Computation: Practice and Experience;2022-08-05