Adoption protocols for fanout-optimal fault-tolerant termination detection-Reference-Cited by-同舟云学术

Adoption protocols for fanout-optimal fault-tolerant termination detection

Published:2013-08-23 Issue:8 Volume:48 Page:13-22
ISSN:0362-1340
Container-title:ACM SIGPLAN Notices
language:en
Short-container-title:SIGPLAN Not.

Author:

Lifflander Jonathan¹,Miller Phil¹,Kale Laxmikant¹

Affiliation:

1. University of Illinois Urbana-Champaign, Urbana, IL, USA

Abstract

Termination detection is relevant for signaling completion (all processors are idle and no messages are in flight) of many operations in distributed systems, including work stealing algorithms, dynamic data exchange, and dynamically structured computations. In the face of growing supercomputers with increasing likelihood that each job may encounter faults, it is important for high-performance computing applications that rely on termination detection that such an algorithm be able to tolerate the inevitable faults. We provide a trio of new practical fault tolerance schemes for a standard approach to termination detection that are easy to implement, present low overhead in both theory and practice, and have scalable costs when recovering from faults. These schemes tolerate all single-process faults, and are probabilistically tolerant of faults affecting multiple processes. We combine the theoretical failure probabilities we can calculate for each algorithm with historical fault records from real machines to show that these algorithms have excellent overall survivability.

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Graphics and Computer-Aided Design,Software

Link

https://dl.acm.org/doi/pdf/10.1145/2517327.2442519

Reference19 articles.

1. Efficient, portable implementation of asynchronous multi-place programs

2. Algorithm-based fault tolerance applied to high performance computing

3. Termination detection for diffusing computations

4. Scalable work stealing

5. Lecture Notes in Computer Science;Freiling F.,2007

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Fault-Tolerant Termination Detection with Safra’s Algorithm;Networked Systems;2021

2. Resilient X10;ACM SIGPLAN Notices;2014-11-26

3. Extreme-Scale Viability of Collective Communication for Resilient Task Scheduling and Work Stealing;2014 44th Annual IEEE/IFIP International Conference on Dependable Systems and Networks;2014-06

4. Coordination Languages and MPI Perturbation Theory: The FOX Tuple Space Framework for Resilience;2014 IEEE International Parallel & Distributed Processing Symposium Workshops;2014-05