Optimizing convolutional neural networks for IoT devices: performance and energy efficiency of quantization techniques-Reference-Cited by-同舟云学术

Optimizing convolutional neural networks for IoT devices: performance and energy efficiency of quantization techniques

Published:2024-02-20 Issue:9 Volume:80 Page:12686-12705
ISSN:0920-8542
Container-title:The Journal of Supercomputing
language:en
Short-container-title:J Supercomput

Author:

Hernández Nicolás,Almeida Francisco,Blanco Vicente

Abstract

AbstractThis document addresses some inherent problems in Machine Learning (ML), such as the high computational and energy costs associated with their implementation on IoT devices. It aims to study and analyze the performance and efficiency of quantization as an optimization method, as well as the possibility of training ML models directly on an IoT device. Quantization involves reducing the precision of model weights and activations while still maintaining acceptable levels of accuracy. Using representative networks for facial recognition developed with TensorFlow and TensorRT, Post-Training Quantization and Quantization-Aware Training are employed to reduce computational load and improve energy efficiency. The computational experience was conducted on a general-purpose computer featuring an Intel i7-1260P processor and an NVIDIA RTX 3080 graphics card used as an accelerator. Additionally, a NVIDIA Jetson AGX Orin was used as an example of an IoT device. We analyze the feasibility of training on an IoT device, the impact of quantization optimization on knowledge transfer-trained models and evaluate the differences between Post-Training Quantization and Quantization-Aware Training in such networks on different devices. Furthermore, the performance and efficiency of NVIDIA’s inference accelerator (Deep Learning Accelerator - DLA, in its 2.0 version) available at the Jetson Orin architecture are studied. We concluded that the Jetson device is capable of performing training on its own. The IoT device can achieve inference performance similar to that of the more powerful processor, thanks to the optimization process, with better energy efficiency. Post-Training Quantization has shown better performance, while Quantization-Aware Training has demonstrated higher energy efficiency. However, since the accelerator cannot execute certain layers of the models, the use of DLA worsens both the performance and efficiency results.

Funder

Ministerio de Ciencia e Innovación

Universidad de la Laguna

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1007/s11227-024-05929-w.pdf

Reference24 articles.

1. Zhang Z, Zhao L, Yang T (2021) Research on the application of artificial intelligence in image recognition technology. J Phys: Conf Ser 1992(3):032118

2. Nassif AB, Shahin I, Attili I, Azzeh M, Shaalan K (2019) Speech recognition using deep neural networks: a systematic review. IEEE Access 7:19143–19165

3. Torfi A, Shirvani RA, Keneshloo Y, Tavaf N, Fox EA (2021) Natural language processing advancements by deep learning: a survey. CoRR. arXiv:abs/2003.01200v4, https://doi.org/10.48550/arXiv.2003.01200

4. Zhuang F, Qi Z, Duan K, Xi D, Zhu Y, Zhu H, Xiong H, He Q (2020) A comprehensive survey on transfer learning. Proc. IEEE 109(1): 43-76. https://doi.org/10.1109/JPROC.2020.3004555

5. Mahapatra S. (2018) Why Deep Learning over Traditional Machine Learning? https://towardsdatascience.com/why-deep-learning-is-needed-over-traditional-machine-learning-1b6a99177063. Accessed 22 Feb 2023