accelerator architectures Latest Research Papers

Analysis of on-chip communication properties in accelerator architectures for deep neural networks

10.1145/3479876.3481588 ◽

2021 ◽

Author(s):

Hana Krichene ◽

Jean-Marc Philippe

Keyword(s):

Neural Networks ◽

Deep Neural Networks ◽

On Chip ◽

Accelerator Architectures

Correction to: Leveraging HPC accelerator architectures with modern techniques — hydrologic modeling on GPUs with ParFlow

Computational Geosciences ◽

10.1007/s10596-021-10065-y ◽

2021 ◽

Author(s):

Jaro Hokkanen ◽

Stefan Kollet ◽

Jiri Kraus ◽

Andreas Herten ◽

Markus Hrywniak ◽

...

Keyword(s):

Hydrologic Modeling ◽

Accelerator Architectures

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication.

10.2172/1808019 ◽

2021 ◽

Author(s):

Gordon Moon ◽

Hyoukjun Kwon ◽

Geonhwa Jeong ◽

prasanth chatarsi ◽

Sivasankaran Rajamanickam ◽

...

Keyword(s):

Matrix Multiplication ◽

Accelerator Architectures

Flynn’s Reconciliation

ACM Transactions on Architecture and Code Optimization ◽

10.1145/3458357 ◽

2021 ◽

Vol 18 (3) ◽

pp. 1-26

Author(s):

Daniel Thuerck ◽

Nicolas Weber ◽

Roberto Bifulco

Keyword(s):

Machine Learning ◽

High Performance Computing ◽

High Performance ◽

General Purpose ◽

Intermediate Representation ◽

Single Source ◽

Architecture Model ◽

Mapping Process ◽

Performance Computing ◽

Accelerator Architectures

A large portion of the recent performance increase in the High Performance Computing (HPC) and Machine Learning (ML) domains is fueled by accelerator cards. Many popular ML frameworks support accelerators by organizing computations as a computational graph over a set of highly optimized, batched general-purpose kernels. While this approach simplifies the kernels’ implementation for each individual accelerator, the increasing heterogeneity among accelerator architectures for HPC complicates the creation of portable and extensible libraries of such kernels. Therefore, using a generalization of the CUDA community’s warp register cache programming idiom, we propose a new programming idiom (CoRe) and a virtual architecture model (PIRCH), abstracting over SIMD and SIMT paradigms. We define and automate the mapping process from a single source to PIRCH’s intermediate representation and develop backends that issue code for three different architectures: Intel AVX512, NVIDIA GPUs, and NEC SX-Aurora. Code generated by our source-to-source compiler for batched kernels, borG, competes favorably with vendor-tuned libraries and is up to 2× faster than hand-tuned kernels across architectures.

Leveraging HPC accelerator architectures with modern techniques — hydrologic modeling on GPUs with ParFlow

Computational Geosciences ◽

10.1007/s10596-021-10051-4 ◽

2021 ◽

Author(s):

Jaro Hokkanen ◽

Stefan Kollet ◽

Jiri Kraus ◽

Andreas Herten ◽

Markus Hrywniak ◽

...

Keyword(s):

Memory Management ◽

High Performance ◽

Hydrologic Model ◽

Third Party ◽

Domain Specific ◽

Scientific Institutions ◽

Significant Performance ◽

Performance Computing ◽

Accelerator Architectures

AbstractRapidly changing heterogeneous supercomputer architectures pose a great challenge to many scientific communities trying to leverage the latest technology in high-performance computing. Many existing projects with a long development history have resultedin a large amount of code that is not directly compatible with the latest accelerator architectures. Furthermore, due to limited resources of scientific institutions, developing and maintaining architecture-specific ports is generally unsustainable. In order to adapt to modern accelerator architectures, many projects rely on directive-based programming models or build the codebase tightly around a third-party domain-specific language or library. This introduces external dependencies out of control of theproject. The presented paper tackles the issue by proposing a lightweight application-side adaptor layer for compute kernels and memory management resulting in a versatile and inexpensive adaptation of new accelerator architectures with little drawbacks.A widely used hydrologic model demonstrates that such an approach pursued more than 20 years ago is still paying off with modern accelerator architectures as demonstrated by a very significant performance gain from NVIDIA A100 GPUs, high developer productivity, and minimally invasive implementation; all while the codebase is kept well maintainable in the long-term.

A survey of accelerator architectures for 3D convolution neural networks

Journal of Systems Architecture ◽

10.1016/j.sysarc.2021.102041 ◽

2021 ◽

pp. 102041

Author(s):

Sparsh Mittal ◽

Vibhu

Keyword(s):

Neural Networks ◽

Convolution Neural Networks ◽

Accelerator Architectures

Performance Analysis of Accelerator Architectures and Programming Models for Parareal Algorithm Solutions of Ordinary Differential Equations

Journal of Computer and Communications ◽

10.4236/jcc.2021.92003 ◽

2021 ◽

Vol 09 (02) ◽

pp. 29-56

Author(s):

Sumathi Lakshmiranganatha ◽

Suresh S. Muknahallipatna

Keyword(s):

Performance Analysis ◽

Differential Equations ◽

Ordinary Differential Equations ◽

Programming Models ◽

Accelerator Architectures

RNN-Based Radio Resource Management on Multicore RISC-V Accelerator Architectures

IEEE Transactions on Very Large Scale Integration (VLSI) Systems ◽

10.1109/tvlsi.2021.3093242 ◽

2021 ◽

pp. 1-14

Author(s):

Gianna Paulin ◽

Renzo Andri ◽

Francesco Conti ◽

Luca Benini

Keyword(s):

Resource Management ◽

Radio Resource Management ◽

Radio Resource ◽

Accelerator Architectures

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication

IEEE Transactions on Parallel and Distributed Systems ◽

10.1109/tpds.2021.3104240 ◽

2021 ◽

pp. 1-1

Author(s):

Gordon Euhyun Moon ◽

Hyoukjun Kwon ◽

Geonhwa Jeong ◽

Prasanth Chatarasi ◽

Sivasankaran Rajamanickam ◽

...

Keyword(s):

Matrix Multiplication ◽

Accelerator Architectures

A Survey of Accelerator Architectures for Deep Neural Networks

Engineering ◽

10.1016/j.eng.2020.01.007 ◽

2020 ◽

Vol 6 (3) ◽

pp. 264-274 ◽

Cited By ~ 14

Author(s):

Yiran Chen ◽

Yuan Xie ◽

Linghao Song ◽

Fan Chen ◽

Tianqi Tang

Keyword(s):

Neural Networks ◽

Deep Neural Networks ◽

Accelerator Architectures

accelerator architectures
Recently Published Documents

TOTAL DOCUMENTS

H-INDEX

Analysis of on-chip communication properties in accelerator architectures for deep neural networks

Correction to: Leveraging HPC accelerator architectures with modern techniques — hydrologic modeling on GPUs with ParFlow

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication.

Flynn’s Reconciliation

Leveraging HPC accelerator architectures with modern techniques — hydrologic modeling on GPUs with ParFlow

A survey of accelerator architectures for 3D convolution neural networks

Performance Analysis of Accelerator Architectures and Programming Models for Parareal Algorithm Solutions of Ordinary Differential Equations

RNN-Based Radio Resource Management on Multicore RISC-V Accelerator Architectures

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication

A Survey of Accelerator Architectures for Deep Neural Networks

Export Citation Format

accelerator architecturesRecently Published Documents

TOTAL DOCUMENTS

H-INDEX

Analysis of on-chip communication properties in accelerator architectures for deep neural networks

Correction to: Leveraging HPC accelerator architectures with modern techniques — hydrologic modeling on GPUs with ParFlow

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication.

Flynn’s Reconciliation

Leveraging HPC accelerator architectures with modern techniques — hydrologic modeling on GPUs with ParFlow

A survey of accelerator architectures for 3D convolution neural networks

Performance Analysis of Accelerator Architectures and Programming Models for Parareal Algorithm Solutions of Ordinary Differential Equations

RNN-Based Radio Resource Management on Multicore RISC-V Accelerator Architectures

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication

A Survey of Accelerator Architectures for Deep Neural Networks

accelerator architectures
Recently Published Documents