Flynn’s Reconciliation-Reference-Cited by-同舟云学术

Flynn’s Reconciliation

Published:2021-06 Issue:3 Volume:18 Page:1-26
ISSN:1544-3566
Container-title:ACM Transactions on Architecture and Code Optimization
language:en
Short-container-title:ACM Trans. Archit. Code Optim.

Author:

Thuerck Daniel¹,Weber Nicolas²,Bifulco Roberto²

Affiliation:

1. NEC Laboratories Europe and TU Darmstadt, Heidelberg, Germany

2. NEC Laboratories Europe, Heidelberg, Germany

Abstract

A large portion of the recent performance increase in the High Performance Computing (HPC) and Machine Learning (ML) domains is fueled by accelerator cards. Many popular ML frameworks support accelerators by organizing computations as a computational graph over a set of highly optimized, batched general-purpose kernels. While this approach simplifies the kernels’ implementation for each individual accelerator, the increasing heterogeneity among accelerator architectures for HPC complicates the creation of portable and extensible libraries of such kernels. Therefore, using a generalization of the CUDA community’s warp register cache programming idiom, we propose a new programming idiom (CoRe) and a virtual architecture model (PIRCH), abstracting over SIMD and SIMT paradigms. We define and automate the mapping process from a single source to PIRCH’s intermediate representation and develop backends that issue code for three different architectures: Intel AVX512, NVIDIA GPUs, and NEC SX-Aurora. Code generated by our source-to-source compiler for batched kernels, borG, competes favorably with vendor-tuned libraries and is up to 2× faster than hand-tuned kernels across architectures.

Publisher

Association for Computing Machinery (ACM)

Subject

Hardware and Architecture,Information Systems,Software

Link

https://dl.acm.org/doi/pdf/10.1145/3458357

Reference56 articles.

1. Automated Compiler Optimization of Multiple Vector Loads/Stores

2. GPU Acceleration of Communication Avoiding Chebyshev Basis Conjugate Gradient Solver for Multiphase CFD Simulations

3. Automatic Vectorization of Interleaved Data Revisited

4. Adaptive precision in block-Jacobi preconditioning for iterative sparse linear system solvers