Providing fault tolerance in extreme scale parallel applications

High-performance computing systems continue to increase in size in the quest for ever higher performance. The resulting increased electronic component count, coupled with the decrease in feature sizes of the silicon manufacturing processes used to build these components, may result in future exascale systems being more susceptible to soft errors caused by cosmic radiation than in current high-performance computing systems. Through the use of techniques such as hardware-based error-correcting codes and checkpoint-restart, many of these faults can be mitigated at the cost of increased hardware overhead, run-time, and energy consumption that can be as much as 10–20%. Some predictions expect these overheads to continue to grow over time. For extreme scale systems, these overheads will represent megawatts of power consumption and millions of dollars of additional hardware costs, which could potentially be avoided with more sophisticated fault-tolerance techniques. In this paper we present new software-based fault tolerance techniques that can be applied to one of the most important classes of software in high-performance computing: iterative sparse matrix solvers. Our new techniques enables us to exploit knowledge of the structure of sparse matrices in such a way as to improve the performance, energy efficiency, and fault tolerance of the overall solution.

Download Full-text

Extreme-scale scripting: Opportunities for large task-parallel applications on petascale computers

Journal of Physics Conference Series ◽

10.1088/1742-6596/180/1/012046 ◽

2009 ◽

Vol 180 ◽

pp. 012046 ◽

Cited By ~ 13

Author(s):

Michael Wilde ◽

Ioan Raicu ◽

Allan Espinosa ◽

Zhao Zhang ◽

Ben Clifford ◽

...

Keyword(s):

Parallel Applications ◽

Task Parallel ◽

Extreme Scale

Download Full-text

Approaches for Parallel Applications Fault Tolerance

Recent Advances in Parallel Virtual Machine and Message Passing Interface - Lecture Notes in Computer Science ◽

10.1007/11846802_2 ◽

2006 ◽

pp. 2-2

Author(s):

Richard L. Graham

Keyword(s):

Fault Tolerance ◽

Parallel Applications

Download Full-text

2018 IEEE/ACM 8th Workshop on Fault Tolerance for HPC at eXtreme Scale (FTXS)

10.1109/ftxs46494.2018 ◽

2018 ◽

Keyword(s):

Fault Tolerance ◽

Extreme Scale

Download Full-text

1st workshop on fault-tolerance for HPC at extreme scale FTXS 2010

2010 International Conference on Dependable Systems and Networks Workshops (DSN-W) ◽

10.1109/dsnw.2010.5542628 ◽

2010 ◽

Author(s):

John Daly ◽

Nathan DeBardeleben

Keyword(s):

Fault Tolerance ◽

Extreme Scale

Download Full-text

HADAB: Enabling Fault Tolerance in Parallel Applications Running in Distributed Environments

Parallel Processing and Applied Mathematics - Lecture Notes in Computer Science ◽

10.1007/978-3-642-31464-3_71 ◽

2012 ◽

pp. 700-709 ◽

Cited By ~ 14

Author(s):

Vania Boccia ◽

Luisa Carracciuolo ◽

Giuliano Laccetti ◽

Marco Lapegna ◽

Valeria Mele

Keyword(s):

Fault Tolerance ◽

Parallel Applications ◽

Distributed Environments

Download Full-text

Parallel Data Partitioning Algorithms for Optimization of Data-Parallel Applications on Modern Extreme-Scale Multicore Platforms for Performance and Energy

IEEE Access ◽

10.1109/access.2018.2879228 ◽

2018 ◽

Vol 6 ◽

pp. 69075-69106 ◽

Cited By ~ 1

Author(s):

Ravi Reddy Manumachu ◽

Alexey Lastovetsky

Keyword(s):

Data Partitioning ◽

Parallel Applications ◽

Multicore Platforms ◽

Data Parallel ◽

Parallel Data ◽

Partitioning Algorithms ◽

Extreme Scale

Download Full-text

Providing fault tolerance through invasive computing

it - Information Technology ◽

10.1515/itit-2016-0022 ◽

2016 ◽

Vol 58 (6) ◽

Cited By ~ 2

Author(s):

Vahid Lari ◽

Andreas Weichslgartner ◽

Alexandru Tanase ◽

Michael Witterauf ◽

Faramarz Khosravi ◽

...

Keyword(s):

Fault Tolerance ◽

Fault Tolerant ◽

Parallel Applications ◽

Technology Scaling ◽

The Face ◽

Tightly Coupled ◽

On Chip ◽

Modular Redundancy ◽

Parallel Arrays ◽

Selection Of

AbstractAs a consequence of technology scaling, today's complex multi-processor systems have become more and more susceptible to errors. In order to satisfy reliability requirements, such systems require methods to detect and tolerate errors. This entails two major challenges: (a) providing a comprehensive approach that ensures fault-tolerant execution of parallel applications across different types of resources, and (b) optimizing resource usage in the face of dynamic fault probabilities or with varying fault tolerance needs of different applications. In this paper, we present a holistic and adaptive approach to provide fault tolerance on Multi-Processor System-on-a-Chip (MPSoC) on demand of an application or environmental needs based on invasive computing. We show how invasive computing may provide adaptive fault tolerance on a heterogeneous MPSoC including hardware accelerators and communication infrastructure such as a Network-on-Chip (NoC). In addition, we present (a) compile-time transformations to automatically adopt well-known redundancy schemes such as Dual Modular Redundancy (DMR) and Triple Modular Redundancy (TMR) for fault-tolerant loop execution on a class of massively parallel arrays of processors called as Tightly Coupled Processor Arrays (). Based on timing characteristics derived from our compilation flow, we further develop (b) a reliability analysis guiding the selection of a suitable degree of fault tolerance. Finally, we present (c) a methodology to detect and adaptively mitigate faults in invasive NoCs.

Download Full-text