Occam: Optimal Data Reuse for Convolutional Neural Networks-Reference-Cited by-同舟云学术

Occam: Optimal Data Reuse for Convolutional Neural Networks

Published:2022-12-16 Issue:1 Volume:20 Page:1-25
ISSN:1544-3566
Container-title:ACM Transactions on Architecture and Code Optimization
language:en
Short-container-title:ACM Trans. Archit. Code Optim.

Author:

Gondimalla Ashish¹^ORCID,Liu Jianqiao²^ORCID,Thottethodi Mithuna¹^ORCID,Vijaykumar T. N.¹^ORCID

Affiliation:

1. Purdue University, West Lafayette, Indiana, USA

2. Google, USA, Mountain View, California, USA

Abstract

Convolutional neural networks (CNNs) are emerging as powerful tools for image processing in important commercial applications. We focus on the important problem of improving the latency of image recognition. While CNNs are highly amenable to prefetching and multithreading to avoid memory latency issues, CNNs’ large data – each layer’s input, filters, and output – poses a memory bandwidth problem. While previous work captures only some of the enormous data reuse, full reuse implies that the initial input image and filters are read once from off-chip and the final output is written once off-chip without spilling the intermediate layers’ data to off-chip. We propose Occam to capture full reuse via four contributions. First, we identify the necessary conditions for full reuse. Second, we identify the dependence closure as the sufficient condition to capture full reuse using the least on-chip memory. Third, because the dependence closure is often too large to fit in on-chip memory, we propose a dynamic programming algorithm that optimally partitions a given CNN to guarantee the least off-chip traffic at the partition boundaries for a given on-chip capacity. While tiling is well-known, our contribution determines the optimal cross-layer tiles. Occam’s partitions reside on different chips, forming a pipeline so that a partition’s filters and dependence closure remain on-chip as different images pass through (i.e., each partition incurs off-chip traffic only for its inputs and outputs). Finally, because the optimal partitions may result in an unbalanced pipeline, we propose staggered asynchronous pipelines (STAPs) that replicate bottleneck stages to improve throughput by staggering mini-batches across replicas. Importantly, STAPs achieve balanced pipelines without changing Occam’s optimal partitioning. Our simulations show that, on average, Occam cuts off-chip transfers by 21× and achieves 2.04× and 1.21× better performance, and 33% better energy than the base case, respectively. Using a field-programmable gate array (FPGA) implementation, Occam performs 6.1× and 1.5× better, on average, than the base case and Layer Fusion, respectively.

Publisher

Association for Computing Machinery (ACM)

Subject

Hardware and Architecture,Information Systems,Software

Link

https://dl.acm.org/doi/pdf/10.1145/3566052

Reference57 articles.

1. 2022. NVIDIA Deep Learning Performance documentation. Retrieved October 11 2022 from https://docs.nvidia.com/deeplearning/performance/dl-performance-convolutional/index.html. Updated May 17 2022.

2. Bit-pragmatic deep neural network computing

3. Cnvlutin: Ineffectual-Neuron-Free Deep Neural Network Computing

4. Manoj Alwani, Han Chen, Michael Ferdman, and Peter Milder. 2016. Fused-layer CNN accelerators. In 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO’16), Taipei, Taiwan. IEEE, 1–12.

5. Analyzing CUDA workloads using a detailed GPU simulator

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Design-Space Exploration of Systolic Array for Edge Inferencing Applications;The Third International Conference on Artificial Intelligence and Machine Learning Systems;2023-10-25