Learning to Learn Gradient Aggregation by Gradient Descent-Reference-Cited by-同舟云学术

Learning to Learn Gradient Aggregation by Gradient Descent

Published:2019-08 Issue: Volume: Page:
ISSN:
Container-title:Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence
language:
Short-container-title:

Author:

Ji Jinlong¹,Chen Xuhui¹²,Wang Qianlong¹,Yu Lixing¹,Li Pan¹

Affiliation:

1. Case Western Reserve University

2. Kent State University

Abstract

In the big data era, distributed machine learning emerges as an important learning paradigm to mine large volumes of data by taking advantage of distributed computing resources. In this work, motivated by learning to learn, we propose a meta-learning approach to coordinate the learning process in the master-slave type of distributed systems. Specifically, we utilize a recurrent neural network (RNN) in the parameter server (the master) to learn to aggregate the gradients from the workers (the slaves). We design a coordinatewise preprocessing and postprocessing method to make the neural network based aggregator more robust. Besides, to address the fault tolerance, especially the Byzantine attack, in distributed machine learning systems, we propose an RNN aggregator with additional loss information (ARNN) to improve the system resilience. We conduct extensive experiments to demonstrate the effectiveness of the RNN aggregator, and also show that it can be easily generalized and achieve remarkable performance when transferred to other distributed systems. Moreover, under majoritarian Byzantine attacks, the ARNN aggregator outperforms the Krum, the state-of-art fault tolerance aggregation method, by 43.14%. In addition, our RNN aggregator enables the server to aggregate gradients from variant local models, which significantly improve the scalability of distributed learning.

Publisher

International Joint Conferences on Artificial Intelligence Organization

Cited by 11 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A meta-active learning approach exploiting instance importance;Expert Systems with Applications;2024-08

2. Overview of Federated Learning and Its Advantages;Advances in Healthcare Information Systems and Administration;2024-04-19

3. Federated Learning with Data-Agnostic Distribution Fusion;2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR);2023-06

4. Reinforcement Learning in Few-Shot Scenarios: A Survey;Journal of Grid Computing;2023-06

5. BiG-FSLF: A Cross Heterogeneous Domain Few-Shot Learning Framework Based on Bidirectional Generation for Hyperspectral Image Change Detection;IEEE Transactions on Geoscience and Remote Sensing;2023