Collaborative Intelligence: Accelerating Deep Neural Network Inference via Device-Edge Synergy-Reference-Cited by-同舟云学术

Collaborative Intelligence: Accelerating Deep Neural Network Inference via Device-Edge Synergy

Published:2020-09-07 Issue: Volume:2020 Page:1-10
ISSN:1939-0114
Container-title:Security and Communication Networks
language:en
Short-container-title:Security and Communication Networks

Author:

Shan Nanliang¹^ORCID,Ye Zecong¹,Cui Xiaolong¹^ORCID

Affiliation:

1. College of Information Engineering, Engineering University of PAP, Xi’an 710086, China

Abstract

With the development of mobile edge computing (MEC), more and more intelligent services and applications based on deep neural networks are deployed on mobile devices to meet the diverse and personalized needs of users. Unfortunately, deploying and inferencing deep learning models on resource-constrained devices are challenging. The traditional cloud-based method usually runs the deep learning model on the cloud server. Since a large amount of input data needs to be transmitted to the server through WAN, it will cause a large service latency. This is unacceptable for most current latency-sensitive and computation-intensive applications. In this paper, we propose Cogent, an execution framework that accelerates deep neural network inference through device-edge synergy. In the Cogent framework, it is divided into two operation stages, including the automatic pruning and partition stage and the containerized deployment stage. Cogent uses reinforcement learning (RL) to automatically predict pruning and partition strategies based on feedback from the hardware configuration and system conditions so that the pruned and partitioned model can better adapt to the system environment and user hardware configuration. Then through containerized deployment to the device and the edge server to accelerate model inference, experiments show that the learning-based hardware-aware automatic pruning and partition scheme can significantly reduce the service latency, and it accelerates the overall model inference process while maintaining accuracy. Using this method can accelerate up to 8.89× without loss of accuracy of more than 7%.

Publisher

Hindawi Limited

Subject

Computer Networks and Communications,Information Systems

Link

http://downloads.hindawi.com/journals/scn/2020/8831341.pdf

Reference25 articles.

1. Amplified locality‐sensitive hashing‐based recommender systems with privacy protection

2. An Online Intrusion Detection System to Cloud Computing Based on Neucube Algorithms

3. Edge Intelligence: Paving the Last Mile of Artificial Intelligence With Edge Computing

4. Distributed Perception by Collaborative Robots

5. Neurosurgeon

Cited by 8 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Researching the CNN Collaborative Inference Mechanism for Heterogeneous Edge Devices;Sensors;2024-06-27

2. Multi-Agent Systems for Collaborative Inference Based on Deep Policy Q-Inference Network;Journal of Grid Computing;2024-02-29

3. Cloud-assisted collaborative inference of convolutional neural networks for vision tasks on resource-constrained devices;Neurocomputing;2023-12

4. An efficient DNN splitting scheme for edge-AI enabled smart manufacturing;Journal of Industrial Information Integration;2023-08

5. A Survey on Collaborative DNN Inference for Edge Intelligence;Machine Intelligence Research;2023-05-03