SparGD: A Sparse GEMM Accelerator with Dynamic Dataflow-Reference-Cited by-同舟云学术

SparGD: A Sparse GEMM Accelerator with Dynamic Dataflow

Published:2024-01-15 Issue:2 Volume:29 Page:1-32
ISSN:1084-4309
Container-title:ACM Transactions on Design Automation of Electronic Systems
language:en
Short-container-title:ACM Trans. Des. Autom. Electron. Syst.

Author:

Wang Bo¹^ORCID,Ma Sheng¹^ORCID,Luo Shengbai¹^ORCID,Wu Lizhou¹^ORCID,Zhang Jianmin¹^ORCID,Zhang Chunyuan¹^ORCID,Li Tiejun¹^ORCID

Affiliation:

1. School of Computer, National University of Defense Technology, China

Abstract

Deep learning has become a highly popular research field, and previously deep learning algorithms ran primarily on CPUs and GPUs. However, with the rapid development of deep learning, it was discovered that existing processors could not meet the specific large-scale computing requirements of deep learning, and custom deep learning accelerators have become popular. The majority of the primary workloads in deep learning are general matrix-matrix multiplications (GEMMs), and emerging GEMMs are highly sparse and irregular. The TPU and SIGMA are typical GEMM accelerators in recent years, but the TPU does not support sparsity, and both the TPU and SIGMA have insufficient utilization rates of the Processing Element (PE). We design and implement SparGD, a sparse GEMM accelerator with dynamic dataflow. SparGD has specific PE structures, flexible distribution networks and reduction networks, and a simple dataflow switching module. When running sparse and irregular GEMMs, SparGD can maintain high PE utilization while utilizing sparsity, and can switch to the optimal dataflow according to the computing environment. For sparse, irregular GEMMs, our experimental results show that SparGD outperforms systolic arrays by 30 times and SIGMA by 3.6 times.

Funder

National Key R&D

NSFC

NSF of Hunan Province

STIP of Hunan Province

Key Laboratory of Advanced Microprocessor Chips and Systems

Publisher

Association for Computing Machinery (ACM)

Subject

Electrical and Electronic Engineering,Computer Graphics and Computer-Aided Design,Computer Science Applications

Link

https://dl.acm.org/doi/pdf/10.1145/3634703

Reference39 articles.

1. Bilge Acun, Matthew Murphy, Xiaodong Wang, Jade Nie, Carole-Jean Wu, and Kim Hazelwood. 2021. Understanding training efficiency of deep learning recommendation models at scale. In 2021 IEEE International Symposium on High-performance Computer Architecture (HPCA’21). IEEE, 802–814.

2. Cnvlutin

3. Shijie Cao, Lingxiao Ma, Wencong Xiao, Chen Zhang, Yunxin Liu, Lintao Zhang, Lanshun Nie, and Zhi Yang. 2019. Seernet: Predicting convolutional neural network feature-map sparsity through low-bit quantization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11216–11225.

4. Amitabha Chakrabarty, Martin Collier, and Sourav Mukhopadhyay. 2009. Matrix-based nonblocking routing algorithm for Beneš networks. In 2009 Computation World: Future Computing, Service Computation, Cognitive, Adaptive, Content, Patterns. IEEE, 551–556.

5. Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Sparm: A Sparse Matrix Multiplication Accelerator Supporting Multiple Dataflows;2024 IEEE 35th International Conference on Application-specific Systems, Architectures and Processors (ASAP);2024-07-24