Fast matrix multiplication via compiler‐only layered data reorganization and intrinsic lowering-Reference-Cited by-同舟云学术

Fast matrix multiplication via compiler‐only layered data reorganization and intrinsic lowering

Published:2023-05-14 Issue:9 Volume:53 Page:1793-1814
ISSN:0038-0644
Container-title:Software: Practice and Experience
language:en
Short-container-title:Softw Pract Exp

Author:

Kuzma Braedy¹,Korostelev Ivan¹,de Carvalho João P. L.¹^ORCID,Moreira José E.²,Barton Christopher³,Araujo Guido⁴,Amaral José Nelson¹

Affiliation:

1. Computing Science Department University of Alberta Edmonton Alberta Canada

2. Thomas J. Watson Research Center IBM Corporation New York New York USA

3. IBM Canada Software Laboratory IBM Corporation Markham Ontario Canada

4. Institute of Computing UNICAMP Campinas São Paulo Brazil

Abstract

AbstractThe resurgence of machine learning has increased the demand for high‐performance basic linear algebra subroutines (BLAS), which have long depended on libraries to achieve peak performance on commodity hardware. High‐performance BLAS implementations rely on a layered approach that consists of tiling and packing layers—for data (re)organization—and micro kernels that perform the actual computations. The algorithm for the tiling and packing layers is target independent but is parameterized to the memory hierarchy and register‐file size. The creation of high‐performance micro kernels requires significant development effort to write tailored assembly code for each architecture. This hand optimization task is complicated by the recent introduction of matrix engines by 's (Matrix Multiply Assist—MMA), (Advanced Matrix eXtensions—AMX), and (Matrix Extensions—ME) to deliver high‐performance matrix operations. This article presents a compiler‐only alternative to the use of high‐performance libraries by incorporating, to the best of our knowledge and for the first time, the automatic generation of the layered approach into LLVM, a production compiler. Modular design of the algorithm, such as the use of LLVM's matrix‐multiply intrinsic for a clear interface between the tiling and packing layers and the micro kernel, makes it easy to retarget the code generation to multiple accelerators. The parameterization of the tiling and packing layers is demonstrated in the generation of code for the MMA unit on IBM's POWER10. This article also describes an algorithm that lowers the matrix‐multiply intrinsic to the MMA unit. The use of intrinsics enables a comprehensive performance study. In processors without hardware matrix engines, the tiling and packing delivers performance up to (Intel)—for small matrices—and more than (POWER9)—for large matrices—faster than PLuTo, a widely used polyhedral optimizer. The performance also approaches high‐performance libraries and is only slower than OpenBLAS and on‐par with Eigen for large matrices. With MMA in POWER10 this solution is, for large matrices, over faster the vector‐extension solution, matches Eigen performance, and achieves up to of BLAS peak performance.

Publisher

Wiley

Subject

Software

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1002/spe.3214

Reference37 articles.

1. VasudevanA AndersonA GreggD.Parallel multi channel convolution using general matrix multiplication. Proceedings of the 2017 IEEE 28th International Conference on Application‐Specific Systems Architectures and Processors (ASAP);2017:19‐24; IEEE.

2. On the Use of BLAS Libraries in Modern Scientific Codes at Scale

3. QinE SamajdarA KwonH et al.SIGMA: A sparse and irregular GEMM accelerator with flexible interconnects for DNN training. Proceedings of the 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA);2020:58‐70.

4. Anatomy of high-performance matrix multiplication

5. XianyiZ QianW YunquanZ.Model‐driven level 3 BLAS performance optimization on Loongson 3A processor. Proceedings of the 2012 IEEE 18th International Conference on Parallel and Distributed Systems;2012:684‐691.