Effects of depth, width, and initialization: A convergence analysis of layer-wise training for deep linear neural networks-Reference-Cited by-同舟云学术

Effects of depth, width, and initialization: A convergence analysis of layer-wise training for deep linear neural networks

Published:2021-12-31 Issue:01 Volume:20 Page:73-119
ISSN:0219-5305
Container-title:Analysis and Applications
language:en
Short-container-title:Anal. Appl.

Author:

Shin Yeonjong¹

Affiliation:

1. Division of Applied Mathematics, Brown University, Providence, RI, 02912, USA

Abstract

Deep neural networks have been used in various machine learning applications and achieved tremendous empirical successes. However, training deep neural networks is a challenging task. Many alternatives have been proposed in place of end-to-end back-propagation. Layer-wise training is one of them, which trains a single layer at a time, rather than trains the whole layers simultaneously. In this paper, we study a layer-wise training using a block coordinate gradient descent (BCGD) for deep linear networks. We establish a general convergence analysis of BCGD and found the optimal learning rate, which results in the fastest decrease in the loss. We identify the effects of depth, width, and initialization. When the orthogonal-like initialization is employed, we show that the width of intermediate layers plays no role in gradient-based training beyond a certain threshold. Besides, we found that the use of deep networks could drastically accelerate convergence when it is compared to those of a depth 1 network, even when the computational cost is considered. Numerical examples are provided to justify our theoretical findings and demonstrate the performance of layer-wise training by BCGD.

Publisher

World Scientific Pub Co Pte Ltd

Subject

Applied Mathematics,Analysis

Link

https://www.worldscientific.com/doi/pdf/10.1142/S0219530521500263

Reference50 articles.

1. Gradient Descent with Identity Initialization Efficiently Learns Positive-Definite Linear Transformations by Deep Residual Networks

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. An Analytic End-to-End Collaborative Deep Learning Algorithm;IEEE Control Systems Letters;2023

2. A Theoretical Framework for End-to-End Learning of Deep Neural Networks With Applications to Robotics;IEEE Access;2023

3. A convergence analysis of Nesterov’s accelerated gradient method in training deep linear neural networks;Information Sciences;2022-10

4. Erratum: Learnability of Quantum Neural Networks [PRX QUANTUM 2 , 040337 (2021)];PRX Quantum;2022-07-11

5. Application of Business Intelligence Based on the Deep Neural Network in Credit Scoring;Security and Communication Networks;2022-05-16