Optimizing Recurrent Neural Networks: A Study on Gradient Normalization of Weights for Enhanced Training Efficiency-Reference-Cited by-同舟云学术

Optimizing Recurrent Neural Networks: A Study on Gradient Normalization of Weights for Enhanced Training Efficiency

Published:2024-07-27 Issue:15 Volume:14 Page:6578
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Wu Xinyi¹,Xiang Bingjie¹,Lu Huaizheng²,Li Chaopeng¹,Huang Xingwang²^ORCID,Huang Weifang¹

Affiliation:

1. School of Ocean Information Engineering, Jimei University, Xiamen 361021, China

2. College of Computer Engineering, Jimei University, Xiamen 361021, China

Abstract

Recurrent Neural Networks (RNNs) are classical models for processing sequential data, demonstrating excellent performance in tasks such as natural language processing and time series prediction. However, during the training of RNNs, the issues of vanishing and exploding gradients often arise, significantly impacting the model’s performance and efficiency. In this paper, we investigate why RNNs are more prone to gradient problems compared to other common sequential networks. To address this issue and enhance network performance, we propose a method for gradient normalization of network weights. This method suppresses the occurrence of gradient problems by altering the statistical properties of RNN weights, thereby improving training effectiveness. Additionally, we analyze the impact of weight gradient normalization on the probability-distribution characteristics of model weights and validate the sensitivity of this method to hyperparameters such as learning rate. The experimental results demonstrate that gradient normalization enhances the stability of model training and reduces the frequency of gradient issues. On the Penn Treebank dataset, this method achieves a perplexity level of 110.89, representing an 11.48% improvement over conventional gradient descent methods. For prediction lengths of 24 and 96 on the ETTm1 dataset, Mean Absolute Error (MAE) values of 0.778 and 0.592 are attained, respectively, resulting in 3.00% and 6.77% improvement over conventional gradient descent methods. Moreover, selected subsets of the UCR dataset show an increase in accuracy ranging from 0.4% to 6.0%. The gradient normalization method enhances the ability of RNNs to learn from sequential and causal data, thereby holding significant implications for optimizing the training effectiveness of RNN-based models.

Funder

National Natural Science Foundation of China

Natural Science Foundation of Xiamen Municipality

Youth Program of the Natural Science Foundation of Fujian Province of China

Publisher

MDPI AG

Link

https://www.mdpi.com/2076-3417/14/15/6578/pdf

Reference27 articles.

1. Survey on recurrent neural network in natural language processing;Tarwani;Int. J. Eng. Trends Technol.,2017

2. Miao, Y., Gowayyed, M., and Metze, F. (2015, January 13–17). EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding. Proceedings of the 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), Scottsdale, AZ, USA.

3. An integrated hybrid CNN–RNN model for visual description and generation of captions;Khamparia;Circuits Syst. Signal Process.,2020

4. Olatunji, I.E., and Cheng, C.H. (2019). Video analytics for visual surveillance and applications: An overview and survey. Machine Learning Paradigms Applications of Learning and Analytics in Intelligent Systems, Springer.

5. Learning long-term dependencies with recurrent neural networks;Schaefer;Neurocomputing,2008