Communication Efficiency and Non-Independent and Identically Distributed Data Challenge in Federated Learning: A Systematic Mapping Study-Reference-Cited by-同舟云学术

Communication Efficiency and Non-Independent and Identically Distributed Data Challenge in Federated Learning: A Systematic Mapping Study

Published:2024-03-24 Issue:7 Volume:14 Page:2720
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Alotaibi Basmah¹²^ORCID,Khan Fakhri Alam¹³⁴^ORCID,Mahmood Sajjad¹³^ORCID

Affiliation:

1. Department of Information and Computer Science, King Fahd University of Petroleum and Minerals, Dhahran 31261, Saudi Arabia

2. Department of Computer Science, College of Computer and Information Sciences, Imam Mohammad Ibn Saud Islamic University (IMSIU), Riyadh 13318, Saudi Arabia

3. Interdisciplinary Research Centre for Intelligent Secure Systems, King Fahd University of Petroleum and Minerals, Dhahran 31261, Saudi Arabia

4. SDAIA-KFUPM Joint Research Center for Artificial Intelligence, King Fahd University of Petroleum and Minerals, Dhahran 31261, Saudi Arabia

Abstract

Federated learning has emerged as a promising approach for collaborative model training across distributed devices. Federated learning faces challenges such as Non-Independent and Identically Distributed (non-IID) data and communication challenges. This study aims to provide in-depth knowledge in the federated learning environment by identifying the most used techniques for overcoming non-IID data challenges and techniques that provide communication-efficient solutions in federated learning. The study highlights the most used non-IID data types, learning models, and datasets in federated learning. A systematic mapping study was performed using six digital libraries, and 193 studies were identified and analyzed after the inclusion and exclusion criteria were applied. We identified that enhancing the aggregation method and clustering are the most widely used techniques for non-IID data problems (used in 18% and 16% of the selected studies), and a quantization technique was the most common technique in studies that provide communication-efficient solutions in federated learning (used in 27% and 15% of the selected studies). Additionally, our work shows that label distribution skew is the most used case to simulate a non-IID environment, specifically, the quantity label imbalance. The supervised learning model CNN model is the most commonly used learning model, and the image datasets MNIST and Cifar-10 are the most widely used datasets when evaluating the proposed approaches. Furthermore, we believe the research community needs to consider the client’s limited resources and the importance of their updates when addressing non-IID and communication challenges to prevent the loss of valuable and unique information. The outcome of this systematic study will benefit federated learning users, researchers, and providers.

Funder

Saudi Data and AI Authority

King Fahd University of Petroleum and Minerals

Publisher

MDPI AG

Link

https://www.mdpi.com/2076-3417/14/7/2720/pdf

Reference239 articles.

1. Federated learning for internet of things: A comprehensive survey;Nguyen;IEEE Commun. Surv. Tutor.,2021

2. A survey of federated learning for edge computing: Research problems and solutions;Xia;High-Confid. Comput.,2021

3. Song, S., and Liang, X. (2024). Federated Pseudo-Sample Clustering Algorithm: A Label-Personalized Federated Learning Scheme Based on Image Clustering. Appl. Sci., 14.

4. Federated multidomain learning with graph ensemble autoencoder GMM for emotion recognition;Zhang;IEEE Trans. Intell. Transp. Syst.,2022

5. Federated learning optimization techniques for non-IID data: A review;Ting;Int. J. Adv. Res. Eng. Technol.,2020