ProKube: Proactive Kubernetes Orchestrator for Inference in Heterogeneous Edge Computing-Reference-Cited by-同舟云学术

ProKube: Proactive Kubernetes Orchestrator for Inference in Heterogeneous Edge Computing

Published:2024-08-18 Issue: Volume: Page:
ISSN:1055-7148
Container-title:International Journal of Network Management
language:en
Short-container-title:Int J Network Mgmt

Author:

Ali Babar¹,Golec Muhammed¹²^ORCID,Singh Gill Sukhpal¹^ORCID,Cuadrado Felix³,Uhlig Steve¹

Affiliation:

1. School of Electronic Engineering and Computer Science Queen Mary University of London London United Kingdom

2. Computer Engineering Department Abdullah Gul University Kayseri Turkey

3. School of Telecommunications Engineering Technical University of Madrid (UPM) Madrid Spain

Abstract

ABSTRACTDeep neural network (DNN) and machine learning (ML) models/ inferences produce highly accurate results demanding enormous computational resources. The limited capacity of end‐user smart gadgets drives companies to exploit computational resources in an edge‐to‐cloud continuum and host applications at user‐facing locations with users requiring fast responses. Kubernetes hosted inferences with poor resource request estimation results in service level agreement (SLA) violation in terms of latency and below par performance with higher end‐to‐end (E2E) delays. Lifetime static resource provisioning either hurts user experience for under‐resource provisioning or incurs cost with over‐provisioning. Dynamic scaling offers to remedy delay by upscaling leading to additional cost whereas a simple migration to another location offering latency in SLA bounds can reduce delay and minimize cost. To address this cost and delay challenges for ML inferences in the inherent heterogeneous, resource‐constrained, and distributed edge environment, we propose ProKube, which is a proactive container scaling and migration orchestrator to dynamically adjust the resources and container locations with a fair balance between cost and delay. ProKube is developed in conjunction with Google Kubernetes Engine (GKE) enabling cross‐cluster migration and/ or dynamic scaling. It further supports the regular addition of freshly collected logs into scheduling decisions to handle unpredictable network behavior. Experiments conducted in heterogeneous edge settings show the efficacy of ProKube to its counterparts cost greedy (CG), latency greedy (LG), and GeKube (GK). ProKube offers 68%, 7%, and 64% SLA violation reduction to CG, LG, and GK, respectively, and it improves cost by 4.77 cores to LG and offers more cost of 3.94 to CG and GK.

Publisher

Wiley

Link

https://onlinelibrary.wiley.com/doi/pdf/10.1002/nem.2298

Reference47 articles.

1. AI‐Based Fog and Edge Computing: A Systematic Review, Taxonomy and Future Directions;Iftikhar S.;Internet of Things,2023

2. Deep Learning in Image Classification Using Residual Network (ResNet) Variants for Detection of Colorectal Cancer;Sarwinda D.;Procedia Computer Science,2021

3. Migration Modeling and Learning Algorithms for Containers in Fog Computing;Tang Z.;IEEE Transactions on Services Computing,2018

4. I.Murturi P. K.Donta andS.Dustdar “Community AI: Towards Community‐Based Federated Learning ” in2023 IEEE 5th International Conference on Cognitive Machine Intelligence (COGMI)(IEEE 2023) 1–9.

5. Y.Hu C.Imes X.Zhao et al. “Pipeline Parallelism for Inference on Heterogeneous Edge Computing ” (2021) arXiv preprint arXiv:2110.14895.