On the Omnipresence of Spurious Local Minima in Certain Neural Network Training Problems-Reference-Cited by-同舟云学术

On the Omnipresence of Spurious Local Minima in Certain Neural Network Training Problems

Published:2023-06-14 Issue: Volume: Page:
ISSN:0176-4276
Container-title:Constructive Approximation
language:en
Short-container-title:Constr Approx

Author:

Christof Constantin,Kowalczyk Julia

Abstract

AbstractWe study the loss landscape of training problems for deep artificial neural networks with a one-dimensional real output whose activation functions contain an affine segment and whose hidden layers have width at least two. It is shown that such problems possess a continuum of spurious (i.e., not globally optimal) local minima for all target functions that are not affine. In contrast to previous works, our analysis covers all sampling and parameterization regimes, general differentiable loss functions, arbitrary continuous nonpolynomial activation functions, and both the finite- and infinite-dimensional setting. It is further shown that the appearance of the spurious local minima in the considered training problems is a direct consequence of the universal approximation theorem and that the underlying mechanisms also cause, e.g.,

$$L^p$$

L p -best approximation problems to be ill-posed in the sense of Hadamard for all networks that do not have a dense image. The latter result also holds without the assumption of local affine linearity and without any conditions on the hidden layers.

Funder

Technische Universität München

Publisher

Springer Science and Business Media LLC

Subject

Computational Mathematics,General Mathematics,Analysis

Link

https://link.springer.com/content/pdf/10.1007/s00365-023-09658-w.pdf

Reference56 articles.

1. Ainsworth, M., Shin, Y.: Plateau phenomenon in gradient descent training of RELU networks: explanation, quantification, and avoidance. SIAM J. Sci. Comput. 43, 3438–3468 (2021)

2. Allen-Zhu, Z., Li, Y., Song, Z.: A convergence theory for deep learning via over-parameterization. In: Chaudhuri, K., Salakhutdinov, R. (eds.) Proceedings of the 36th International Conference on Machine Learning, vol. 97, pp. 242–252, PMLR (2019)

3. Arjevani, Y., Field, M.: Analytic study of families of spurious minima in two-layer ReLU neural networks: a tale of symmetry II. In: Advances in Neural Information Processing Systems, vol. 34. Curran Associates, Inc. (2021)

4. Auer, P., Herbster, M., Warmuth, M.K.: Exponentially many local minima for single neurons. In: Touretzky, D.S., Mozer, M.C., Hasselmo, M.E. (eds.) Advances in Neural Information Processing Systems, vol. 8, pp. 316–322. Curran Associates, Inc. (1996)

5. Benedetto, J.J., Czaja, W.: Integration and Modern Analysis. Birkhäuser Advanced Texts. Birkhäuser, Boston (2010)

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. On the identification and optimization of nonsmooth superposition operators in semilinear elliptic PDEs;ESAIM: Control, Optimisation and Calculus of Variations;2023-12-19