GPU-accelerated molecular dynamics: State-of-art software performance and porting from Nvidia CUDA to AMD HIP

The International Journal of High Performance Computing Applications ◽

10.1177/10943420211008288 ◽

2021 ◽

pp. 109434202110082

Author(s):

Nikolay Kondratyuk ◽

Vsevolod Nikolskiy ◽

Daniil Pavlov ◽

Vladimir Stegailov

Keyword(s):

Molecular Dynamics ◽

High Performance ◽

Software Performance ◽

Computing Systems ◽

Accelerated Molecular Dynamics ◽

Nvidia Cuda ◽

Software And Hardware ◽

Management Capabilities ◽

Utilization Time ◽

Performance Computing

Classical molecular dynamics (MD) calculations represent a significant part of the utilization time of high-performance computing systems. As usual, the efficiency of such calculations is based on an interplay of software and hardware that are nowadays moving to hybrid GPU-based technologies. Several well-developed open-source MD codes focused on GPUs differ both in their data management capabilities and in performance. In this work, we analyze the performance of LAMMPS, GROMACS and OpenMM MD packages with different GPU backends on Nvidia Volta and AMD Vega20 GPUs. We consider the efficiency of solving two identical MD models (generic for material science and biomolecular studies) using different software and hardware combinations. We describe our experience in porting the CUDA backend of LAMMPS to ROCm HIP that shows considerable benefits for AMD GPUs comparatively to the OpenCL backend.

MonSTer: An Out-of-the-Box Monitoring Tool for High Performance Computing Systems

2020 IEEE International Conference on Cluster Computing (CLUSTER) ◽

10.1109/cluster49012.2020.00022 ◽

2020 ◽

Author(s):

Jie Li ◽

Ghazanfar Ali ◽

Ngan Nguyen ◽

Jon Hass ◽

Alan Sill ◽

...

Keyword(s):

High Performance Computing ◽

High Performance ◽

Monitoring Tool ◽

Computing Systems ◽

Performance Computing

Session details: Special issue on the 1st international workshop on performance modeling, benchmarking and simulation of high performance computing systems (PMBS 10)

ACM SIGMETRICS Performance Evaluation Review ◽

10.1145/3263957 ◽

2011 ◽

Vol 38 (4) ◽

Keyword(s):

High Performance Computing ◽

High Performance ◽

Performance Modeling ◽

International Workshop ◽

Special Issue ◽

Computing Systems ◽

Performance Computing

Treasure Hunt Framework: Distributing Metaheuristics on High Performance Computing Systems

Swarm and Evolutionary Computation ◽

10.1016/j.swevo.2021.100906 ◽

2021 ◽

pp. 100906

Author(s):

Peter Frank Perroni ◽

Myriam Regattieri Delgado ◽

Daniel Weingaertner

Keyword(s):

High Performance Computing ◽

High Performance ◽

Computing Systems ◽

Performance Computing

Session details: Special issue on the 2nd international workshop on performance modeling, benchmarking and simulation of high performance computing systems (PMBS 11)

ACM SIGMETRICS Performance Evaluation Review ◽

10.1145/3264251 ◽

2012 ◽

Vol 40 (2) ◽

Keyword(s):

High Performance Computing ◽

High Performance ◽

Performance Modeling ◽

International Workshop ◽

Special Issue ◽

Computing Systems ◽

Performance Computing

Achieving Safety for Power Shifting in Overprovisioned High Performance Computing Systems

2016 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW) ◽

10.1109/ipdpsw.2016.160 ◽

2016 ◽

Author(s):

Shirley Moore

Keyword(s):

High Performance Computing ◽

High Performance ◽

Computing Systems ◽

Performance Computing

Overview of AIM: supporting computer vision on heterogeneous high-performance computing systems

10.1117/12.323476 ◽

1998 ◽

Author(s):

Monica Sweat ◽

Joseph N. Wilson

Keyword(s):

Computer Vision ◽

High Performance Computing ◽

High Performance ◽

Computing Systems ◽

Performance Computing

A Novel Energy Efficient Scheduling for High Performance Computing Systems

2018 9th International Conference on Computing, Communication and Networking Technologies (ICCCNT) ◽

10.1109/icccnt.2018.8494120 ◽

2018 ◽

Cited By ~ 4

Author(s):

Tarun Biswas ◽

Pratyay Kuila ◽

Anjan Kumar Ray

Keyword(s):

High Performance Computing ◽

Energy Efficient ◽

High Performance ◽

Computing Systems ◽

Performance Computing ◽

Energy Efficient Scheduling

Merging Plasmonics and Silicon Photonics Towards Greener and Faster “Network-on-Chip” Solutions for Data Centers and High-Performance Computing Systems

Plasmonics - Principles and Applications ◽

10.5772/51853 ◽

2012 ◽

Cited By ~ 3

Author(s):

Sotirios Papaioannou ◽

Konstantinos Vyrsokinos ◽

Dimitrios Kalavrouziotis ◽

Giannis Giannoulis ◽

Dimitrios Apostolopoulos ◽

...

Keyword(s):

High Performance Computing ◽

Silicon Photonics ◽

High Performance ◽

Data Centers ◽

Network On Chip ◽

Computing Systems ◽

On Chip ◽

Performance Computing

Application-based fault tolerance techniques for sparse matrix solvers

The International Journal of High Performance Computing Applications ◽

10.1177/1094342017694946 ◽

2017 ◽

Vol 32 (5) ◽

pp. 627-640

Author(s):

Simon McIntosh–Smith ◽

Rob Hunt ◽

James Price ◽

Alex Warwick Vesztrocy

Keyword(s):

Fault Tolerance ◽

High Performance Computing ◽

High Performance ◽

Sparse Matrix ◽

Sparse Matrices ◽

Error Correcting Codes ◽

Computing Systems ◽

Hardware Costs ◽

Extreme Scale ◽

Performance Computing

High-performance computing systems continue to increase in size in the quest for ever higher performance. The resulting increased electronic component count, coupled with the decrease in feature sizes of the silicon manufacturing processes used to build these components, may result in future exascale systems being more susceptible to soft errors caused by cosmic radiation than in current high-performance computing systems. Through the use of techniques such as hardware-based error-correcting codes and checkpoint-restart, many of these faults can be mitigated at the cost of increased hardware overhead, run-time, and energy consumption that can be as much as 10–20%. Some predictions expect these overheads to continue to grow over time. For extreme scale systems, these overheads will represent megawatts of power consumption and millions of dollars of additional hardware costs, which could potentially be avoided with more sophisticated fault-tolerance techniques. In this paper we present new software-based fault tolerance techniques that can be applied to one of the most important classes of software in high-performance computing: iterative sparse matrix solvers. Our new techniques enables us to exploit knowledge of the structure of sparse matrices in such a way as to improve the performance, energy efficiency, and fault tolerance of the overall solution.

Predicting Node Failure in High Performance Computing Systems from Failure and Usage Logs

2011 IEEE International Symposium on Parallel and Distributed Processing Workshops and Phd Forum ◽

10.1109/ipdps.2011.310 ◽

2011 ◽

Cited By ~ 15

Author(s):

Nithin Nakka ◽

Ankit Agrawal ◽

Alok Choudhary

Keyword(s):

High Performance Computing ◽

High Performance ◽

Computing Systems ◽

Node Failure ◽

Performance Computing