Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions-Reference-Cited by-同舟云学术

Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions

Published:2023-12-14 Issue:4 Volume:20 Page:1-26
ISSN:1544-3566
Container-title:ACM Transactions on Architecture and Code Optimization
language:en
Short-container-title:ACM Trans. Archit. Code Optim.

Author:

Ferrari Victor¹^ORCID,Sousa Rafael¹^ORCID,Pereira Marcio¹^ORCID,L. De Carvalho João P.²^ORCID,Amaral José Nelson²^ORCID,Moreira José³^ORCID,Araujo Guido⁴^ORCID

Affiliation:

1. Institute of Computing-UNICAMP, Brazil

2. University of Alberta, Canada

3. IBM Research, United States of America

4. Institute of Computing–UNICAMP, Brazil

Abstract

Convolution is one of the most computationally intensive operations that must be performed for machine learning model inference. A traditional approach to computing convolutions is known as the Im2Col + BLAS method. This article proposes SConv: a direct-convolution algorithm based on an MLIR/LLVM code-generation toolchain that can be integrated into machine-learning compilers. This algorithm introduces: (a) Convolution Slicing Analysis (CSA)—a convolution-specific 3D cache-blocking analysis pass that focuses on tile reuse over the cache hierarchy; (b) Convolution Slicing Optimization—a code-generation pass that uses CSA to generate a tiled direct-convolution macro-kernel; and (c) Vector-based Packing—an architecture-specific optimized input-tensor packing solution based on vector-register shift instructions for convolutions with unitary stride. Experiments conducted on 393 convolutions from full ONNX-MLIR machine learning models indicate that the elimination of the Im2Col transformation and the use of fast packing routines result in a total packing time reduction, on full model inference, of 2.3×–4.0× on Intel x86 and 3.3×–5.9× on IBM POWER10. The speed-up over an Im2Col + BLAS method based on current BLAS implementations for end-to-end machine-learning model inference is in the range of 11%–27% for Intel x86 and 11%–34% for IBM POWER10 architectures. The total convolution speedup for model inference is 13%–28% on Intel x86 and 23%–39% on IBM POWER10. SConv also outperforms BLAS GEMM, when computing pointwise convolutions in more than 82% of the 219 tested instances.

Publisher

Association for Computing Machinery (ACM)

Subject

Hardware and Architecture,Information Systems,Software

Link

https://dl.acm.org/doi/pdf/10.1145/3625004

Reference36 articles.

1. High-Performance Low-Memory Lowering: GEMM-based Algorithms for DNN Convolution

2. Reformulating the direct convolution for high-performance deep learning inference on ARM processors

3. Compiling for the IBM Matrix Engine for Enterprise Workloads

4. Kumar Chellapilla, Sidd Puri, and Patrice Simard. 2006. High performance convolutional neural networks for document processing. In Proceedings of the 10th International Workshop on Frontiers in Handwriting Recognition, Guy Lorette (Ed.). Université de Rennes 1, Suvisoft, La Baule (France). Retrieved from https://hal.inria.fr/inria-00112631

5. Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An automated end-to-end optimizing compiler for deep learning. In Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation (OSDI’18). USENIX Association, 579–594.

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Improving Direct Convolution through Tensor Slicing, Vectorized Packing and ISA Extensions;Anais do XXXVII Concurso de Teses e Dissertações (CTD 2024);2024-07-21

2. SIMD-Constrained Lookup Table for Accelerating Variable-Weighted Convolution on x86/64 CPUs;IEEE Access;2024