Application-based fault tolerance techniques for sparse matrix solvers-Reference-Cited by-同舟云学术

Application-based fault tolerance techniques for sparse matrix solvers

Published:2017-05-10 Issue:5 Volume:32 Page:627-640
ISSN:1094-3420
Container-title:The International Journal of High Performance Computing Applications
language:en
Short-container-title:The International Journal of High Performance Computing Applications

Author:

McIntosh–Smith Simon¹,Hunt Rob¹,Price James¹,Vesztrocy Alex Warwick²

Affiliation:

1. University of Bristol, UK

2. University College London, UK

Abstract

High-performance computing systems continue to increase in size in the quest for ever higher performance. The resulting increased electronic component count, coupled with the decrease in feature sizes of the silicon manufacturing processes used to build these components, may result in future exascale systems being more susceptible to soft errors caused by cosmic radiation than in current high-performance computing systems. Through the use of techniques such as hardware-based error-correcting codes and checkpoint-restart, many of these faults can be mitigated at the cost of increased hardware overhead, run-time, and energy consumption that can be as much as 10–20%. Some predictions expect these overheads to continue to grow over time. For extreme scale systems, these overheads will represent megawatts of power consumption and millions of dollars of additional hardware costs, which could potentially be avoided with more sophisticated fault-tolerance techniques. In this paper we present new software-based fault tolerance techniques that can be applied to one of the most important classes of software in high-performance computing: iterative sparse matrix solvers. Our new techniques enables us to exploit knowledge of the structure of sparse matrices in such a way as to improve the performance, energy efficiency, and fault tolerance of the overall solution.

Funder

Seventh Framework Programme

Engineering and Physical Sciences Research Council

Publisher

SAGE Publications

Subject

Hardware and Architecture,Theoretical Computer Science,Software

Link

http://journals.sagepub.com/doi/pdf/10.1177/1094342017694946

Reference20 articles.

1. Unprotected Computing: A Large-Scale Study of DRAM Raw Error Rate on a Supercomputer

2. Comparison of accelerated DRAM soft error rates measured at component and system level

3. The university of Florida sparse matrix collection

4. The International Exascale Software Project roadmap

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Evaluating the Resiliency of Posits for Scientific Computing;Proceedings of the SC '23 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis;2023-11-12