An algorithm/hardware co‐optimized method to accelerate CNNs with compressed convolutional weights on FPGA-Reference-Cited by-同舟云学术

An algorithm/hardware co‐optimized method to accelerate CNNs with compressed convolutional weights on FPGA

Published:2024-01-06 Issue:11 Volume:36 Page:
ISSN:1532-0626
Container-title:Concurrency and Computation: Practice and Experience
language:en
Short-container-title:Concurrency and Computation

Author:

Shang Jiangwei¹^ORCID,Zhang Zhan¹,Zhang Kun²,Li Chuanyou³,Qian Lei²,Liu Hongwei¹

Affiliation:

1. School of Computer Science and Technology Harbin Institute of Technology Harbin China

2. State Key Laboratory of Mathematical Engineering and Advanced Computing Wuxi China

3. School of Computer Science and Engineering Southeast University Nanjing China

Abstract

SummaryConvolutional neural networks (CNNs) have shown remarkable advantages in a wide range of domains at the expense of huge parameters and computations. Modern CNNs still tend to be more complex and larger to achieve better inference accuracy. However, the complex and large structures of CNNs could slow down the inference speed. Recently, Compressing the convolutional weights to be sparse by pruning the unimportant parameters has been demonstrated as an efficient way to reduce the computations of CNNs. On the other hand, field‐programmable gate arrays (FPGAs) have been a popular hardware platform to accelerate CNN inference. In this paper, we propose an algorithm/hardware co‐optimized method for accelerating CNN inference on FPGAs. For the algorithm, we take advantage of unstructured and structured parameter sparsifying methods to achieve high sparsity and keep the regularity of convolutional weights. Correspondingly, hardware‐friendly index representations of sparse convolutional weights are proposed. For the hardware architecture, we propose row‐wise input‐stationary dataflow, which is tightly coupled with the algorithm. A row‐wise computing engine (RConv Engine) is proposed, which is based on the dataflow. Inside the RConv Engine, the scalar‐vector structure is applied to implement the basic processing elements (PEs). To flexibly calculate the feature map with various sizes, the PEs are organized in a 2D structure with two work modes. The experimental results demonstrate that our co‐optimized method implements high sparsity of convolutional weights, and the computing engine achieves high computation efficiency. Compared with other accelerators, our co‐optimized method implements a 10.9 speedup on FPS at most with the highest sparsity of convolutional weights and negligible accuracy loss.

Funder

National Natural Science Foundation of China

Publisher

Wiley

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1002/cpe.8011

Reference27 articles.

1. SimonyanK ZissermanA.Very deep convolutional networks for large‐scale image recognition. arXiv preprint arXiv:1409.1556.2014.

2. RedmonJ DivvalaS GirshickR FarhadiA.You only look once: unified real‐time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE.2016.

3. HeK ZhangX RenS SunJ.Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE.2016.

4. Angel-Eye: A Complete Design Flow for Mapping CNN Onto Embedded FPGA

5. ZengH ChenR ZhangC PrasannaV.A framework for generating high throughput CNN implementations on FPGAs. Proceedings of the 2018 ACM/SIGDA International Symposium on Field‐Programmable Gate Arrays. ACM.2018:117‐126.