An Empirical Study of the Impact of Hyperparameter Tuning and Model Optimization on the Performance Properties of Deep Neural Networks-Reference-Cited by-同舟云学术

An Empirical Study of the Impact of Hyperparameter Tuning and Model Optimization on the Performance Properties of Deep Neural Networks

Published:2022-04-09 Issue:3 Volume:31 Page:1-40
ISSN:1049-331X
Container-title:ACM Transactions on Software Engineering and Methodology
language:en
Short-container-title:ACM Trans. Softw. Eng. Methodol.

Author:

Liao Lizhi¹^ORCID,Li Heng²^ORCID,Shang Weiyi¹,Ma Lei³

Affiliation:

1. Concordia University, Montréal, Québec, Canada

2. Polytechnique Montréal, Montréal, Québec, Canada

3. University of Alberta, Edmonton, Alberta, Canada

Abstract

Deep neural network (DNN) models typically have many hyperparameters that can be configured to achieve optimal performance on a particular dataset. Practitioners usually tune the hyperparameters of their DNN models by training a number of trial models with different configurations of the hyperparameters, to find the optimal hyperparameter configuration that maximizes the training accuracy or minimizes the training loss. As such hyperparameter tuning usually focuses on the model accuracy or the loss function, it is not clear and remains under-explored how the process impacts other performance properties of DNN models, such as inference latency and model size. On the other hand, standard DNN models are often large in size and computing-intensive, prohibiting them from being directly deployed in resource-bounded environments such as mobile devices and Internet of Things (IoT) devices. To tackle this problem, various model optimization techniques (e.g., pruning or quantization) are proposed to make DNN models smaller and less computing-intensive so that they are better suited for resource-bounded environments. However, it is neither clear how the model optimization techniques impact other performance properties of DNN models such as inference latency and battery consumption, nor how the model optimization techniques impact the effect of hyperparameter tuning (i.e., the compounding effect). Therefore, in this paper, we perform a comprehensive study on four representative and widely-adopted DNN models, i.e., CNN image classification , Resnet-50 , CNN text classification , and LSTM sentiment classification , to investigate how different DNN model hyperparameters affect the standard DNN models, as well as how the hyperparameter tuning combined with model optimization affect the optimized DNN models, in terms of various performance properties (e.g., inference latency or battery consumption). Our empirical results indicate that tuning specific hyperparameters has heterogeneous impact on the performance of DNN models across different models and different performance properties. In particular, although the top tuned DNN models usually have very similar accuracy, they may have significantly different performance in terms of other aspects (e.g., inference latency). We also observe that model optimization has a confounding effect on the impact of hyperparameters on DNN model performance. For example, two sets of hyperparameters may result in standard models with similar performance but their performance may become significantly different after they are optimized and deployed on the mobile device. Our findings highlight that practitioners can benefit from paying attention to a variety of performance properties and the confounding effect of model optimization when tuning and optimizing their DNN models.

Publisher

Association for Computing Machinery (ACM)

Subject

Software

Link

https://dl.acm.org/doi/pdf/10.1145/3506695

Reference94 articles.

1. https://towardsdatascience.com/hyperparameters-in-deep-learning-927f7b2084dd 2019 Hyperparameters in Deep Learning

2. https://github.com/keras-team/keras-io/tree/master/examples 2020 Keras Code Examples

3. https://www.tensorflow.org/model_optimization/guide/pruning/pruning_with_keras 2020 Pruning in Keras Example

4. https://www.tensorflow.org/model_optimization 2020 Tensorflow Model Optimization

5. Martín Abadi Ashish Agarwal Paul Barham Eugene Brevdo Yuan Yu et al.2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. https://www.tensorflow.org/. Software available from tensorflow.org.

Cited by 49 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. ECGencode: Compact and computationally efficient deep learning feature encoder for ECG signals;Expert Systems with Applications;2024-12

2. CDK: A novel high-performance transfer feature technique for early detection of osteoarthritis;Journal of Pathology Informatics;2024-12

3. Multi-output neural network model for predicting biochar yield and composition;Science of The Total Environment;2024-10

4. Globalizing Food Items Based on Ingredient Consumption;Sustainability;2024-08-30

5. Optimizing CNN-LSTM for the Localization of False Data Injection Attacks in Power Systems;Applied Sciences;2024-08-06