The effective noise of stochastic gradient descent-Reference-Cited by-同舟云学术

The effective noise of stochastic gradient descent

Published:2022-08-01 Issue:8 Volume:2022 Page:083405
ISSN:1742-5468
Container-title:Journal of Statistical Mechanics: Theory and Experiment
language:
Short-container-title:J. Stat. Mech.

Author:

Mignacco Francesca,Urbani Pierfrancesco

Abstract

Abstract Stochastic gradient descent (SGD) is the workhorse algorithm of deep learning technology. At each step of the training phase, a mini batch of samples is drawn from the training dataset and the weights of the neural network are adjusted according to the performance on this specific subset of examples. The mini-batch sampling procedure introduces a stochastic dynamics to the gradient descent, with a non-trivial state-dependent noise. We characterize the stochasticity of SGD and a recently-introduced variant, persistent SGD, in a prototypical neural network model. In the under-parametrized regime, where the final training error is positive, the SGD dynamics reaches a stationary state and we define an effective temperature from the fluctuation–dissipation theorem, computed from dynamical mean-field theory. We use the effective temperature to quantify the magnitude of the SGD noise as a function of the problem parameters. In the over-parametrized regime, where the training error vanishes, we measure the noise magnitude of SGD by computing the average distance between two replicas of the system with the same initialization and two different realizations of SGD noise. We find that the two noise measures behave similarly as a function of the problem parameters. Moreover, we observe that noisier algorithms lead to wider decision boundaries of the corresponding constraint satisfaction problem.

Publisher

IOP Publishing

Subject

Statistics, Probability and Uncertainty,Statistics and Probability,Statistical and Nonlinear Physics

Link

https://iopscience.iop.org/article/10.1088/1742-5468/ac841d/pdf

Reference59 articles.

1. Deep learning

2. Understanding deep learning is also a job for physicists

3. Representation Learning: A Review and New Perspectives

4. Tencent ML-Images: A Large-Scale Multi-Label Image Database for Visual Representation Learning

5. Online learning and stochastic approximations;Bottou,1998

Cited by 18 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Weight fluctuations in deep linear neural networks and a derivation of the inverse-variance flatness relation;Physical Review Research;2024-07-25

2. SEVEN: Pruning Transformer Model by Reserving Sentinels;2024 International Joint Conference on Neural Networks (IJCNN);2024-06-30

3. Stochastic Gradient Descent-like relaxation is equivalent to Metropolis dynamics in discrete optimization and inference problems;Scientific Reports;2024-05-21

4. Rigorous Dynamical Mean-Field Theory for Stochastic Gradient Descent Methods;SIAM Journal on Mathematics of Data Science;2024-05-06

5. On the different regimes of stochastic gradient descent;Proceedings of the National Academy of Sciences;2024-02-20