Optimizing Deep Learning Inference on Embedded Systems Through Adaptive Model Selection-Reference-Cited by-同舟云学术

Optimizing Deep Learning Inference on Embedded Systems Through Adaptive Model Selection

Published:2020-01-31 Issue:1 Volume:19 Page:1-28
ISSN:1539-9087
Container-title:ACM Transactions on Embedded Computing Systems
language:en
Short-container-title:ACM Trans. Embed. Comput. Syst.

Author:

Marco Vicent Sanz¹,Taylor Ben²^ORCID,Wang Zheng³,Elkhatib Yehia²

Affiliation:

1. Osaka University, Japan

2. Lancaster University, United Kingdom

3. University of Leeds, United Kingdom

Abstract

Deep neural networks (DNNs) are becoming a key enabling technique for many application domains. However, on-device inference on battery-powered, resource-constrained embedding systems is often infeasible due to prohibitively long inferencing time and resource requirements of many DNNs. Offloading computation into the cloud is often unacceptable due to privacy concerns, high latency, or the lack of connectivity. Although compression algorithms often succeed in reducing inferencing times, they come at the cost of reduced accuracy. This article presents a new, alternative approach to enable efficient execution of DNNs on embedded devices. Our approach dynamically determines which DNN to use for a given input by considering the desired accuracy and inference time. It employs machine learning to develop a low-cost predictive model to quickly select a pre-trained DNN to use for a given input and the optimization constraint. We achieve this first by offline training a predictive model and then using the learned model to select a DNN model to use for new, unseen inputs. We apply our approach to two representative DNN domains: image classification and machine translation. We evaluate our approach on a Jetson TX2 embedded deep learning platform and consider a range of influential DNN models including convolutional and recurrent neural networks. For image classification, we achieve a 1.8x reduction in inference time with a 7.52% improvement in accuracy over the most capable single DNN model. For machine translation, we achieve a 1.34x reduction in inference time over the most capable single model with little impact on the quality of translation.

Funder

Engineering and Physical Sciences Research Council

Publisher

Association for Computing Machinery (ACM)

Subject

Hardware and Architecture,Software

Link

https://dl.acm.org/doi/pdf/10.1145/3371154

Reference81 articles.

1. J. J. Allaire Dirk Eddelbuettel Nick Golding and Yuan Tang. 2016. TensorFlow for R. Available at https://tensorflow.rstudio.com/. J. J. Allaire Dirk Eddelbuettel Nick Golding and Yuan Tang. 2016. TensorFlow for R. Available at https://tensorflow.rstudio.com/.

2. Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv:1409.0473. Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv:1409.0473.

3. Jiawang Bai Yiming Li Jiawei Li Yong Jiang and Shutao Xia. 2019. Rectified decision trees: Towards interpretability compression and empirical soundness. arxiv:1903.05965. Jiawang Bai Yiming Li Jiawei Li Yong Jiang and Shutao Xia. 2019. Rectified decision trees: Towards interpretability compression and empirical soundness. arxiv:1903.05965.

Cited by 43 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Certified Quantization Strategy Synthesis for Neural Networks;Lecture Notes in Computer Science;2024-09-11

2. Adapting Neural Networks at Runtime: Current Trends in At-Runtime Optimizations for Deep Learning;ACM Computing Surveys;2024-05-14

3. Artificial intelligence and edge computing for machine maintenance-review;Artificial Intelligence Review;2024-04-15

4. Green Edge AI: A Contemporary Survey;Proceedings of the IEEE;2024

5. ECADA: An Edge Computing Assisted Delay-Aware Anomaly Detection Scheme for ICS;2023 19th International Conference on Mobility, Sensing and Networking (MSN);2023-12-14