Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks-Reference-Cited by-同舟云学术

Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks

Published:2020-07 Issue: Volume: Page:
ISSN:
Container-title:Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence
language:
Short-container-title:

Author:

Chen Jinghui¹,Zhou Dongruo¹,Tang Yiqi²,Yang Ziyan³,Cao Yuan¹,Gu Quanquan¹

Affiliation:

1. University of California, Los Angeles

2. Ohio State University

3. University of Virginia

Abstract

Adaptive gradient methods, which adopt historical gradient information to automatically adjust the learning rate, despite the nice property of fast convergence, have been observed to generalize worse than stochastic gradient descent (SGD) with momentum in training deep neural networks. This leaves how to close the generalization gap of adaptive gradient methods an open problem. In this work, we show that adaptive gradient methods such as Adam, Amsgrad, are sometimes "over adapted". We design a new algorithm, called Partially adaptive momentum estimation method, which unifies the Adam/Amsgrad with SGD by introducing a partial adaptive parameter $p$, to achieve the best from both worlds. We also prove the convergence rate of our proposed algorithm to a stationary point in the stochastic nonconvex optimization setting. Experiments on standard benchmarks show that our proposed algorithm can maintain fast convergence rate as Adam/Amsgrad while generalizing as well as SGD in training deep neural networks. These results would suggest practitioners pick up adaptive gradient methods once again for faster training of deep neural networks.

Publisher

International Joint Conferences on Artificial Intelligence Organization

Cited by 12 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Channel estimation for RIS-aided MIMO systems in MmWave wireless communications with a few active elements;Cluster Computing;2024-07-17

2. Communication-Efficient Zeroth-Order Adaptive Optimization for Federated Learning;Mathematics;2024-04-11

3. Flatness-Aware Minimization for Domain Generalization;2023 IEEE/CVF International Conference on Computer Vision (ICCV);2023-10-01

4. Theoretical analysis of Adam using hyperparameters close to one without Lipschitz smoothness;Numerical Algorithms;2023-07-04

5. Towards Faster Training Algorithms Exploiting Bandit Sampling From Convex to Strongly Convex Conditions;IEEE Transactions on Emerging Topics in Computational Intelligence;2023-04