Scheduling of Time-Varying Workloads Using Reinforcement Learning-Reference-Cited by-同舟云学术

Scheduling of Time-Varying Workloads Using Reinforcement Learning

Published:2021-05-18 Issue:10 Volume:35 Page:9000-9008
ISSN:2374-3468
Container-title:Proceedings of the AAAI Conference on Artificial Intelligence
language:
Short-container-title:AAAI

Author:

Mondal Shanka Subhra,Sheoran Nikhil,Mitra Subrata

Abstract

Resource usage of production workloads running on shared compute clusters often fluctuate significantly across time. While simultaneous spike in the resource usage between two workloads running on the same machine can create performance degradation, unused resources in a machine results in wastage and undesirable operational characteristics for a compute cluster. Prior works did not consider such temporal resource fluctuations or their alignment for scheduling decisions. Due to the variety of time-varying workloads, their complex resource usage characteristics, it is challenging to design well-defined heuristics for scheduling them optimally across different machines in a cluster. In this paper, we propose a Deep Reinforcement Learning (DRL) based approach to exploit various temporal resource usage patterns of time varying workloads as well as a technique for creating equivalence classes among a large number of production workloads to improve scalability of our method. Validations with real production traces from Google and Alibaba show that our technique can significantly improve metrics for operational excellence (e.g. utilization, fragmentation, resource exhaustion etc.) for a cluster, compared to the baselines.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Subject

General Medicine

Cited by 11 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. PheCon: Fine-Grained VM Consolidation with Nimble Resource Defragmentation in Public Cloud Platforms;Proceedings of the 53rd International Conference on Parallel Processing;2024-08-12

2. BCEdge: SLO-Aware DNN Inference Services With Adaptive Batch-Concurrent Scheduling on Edge Devices;IEEE Transactions on Network and Service Management;2024-08

3. Batch Jobs Load Balancing Scheduling in Cloud Computing Using Distributional Reinforcement Learning;IEEE Transactions on Parallel and Distributed Systems;2024-01

4. Gödel;Proceedings of the 2023 ACM Symposium on Cloud Computing;2023-10-30

5. Action Masked Deep Reinforcement learning for Controlling Industrial Assembly Lines;2023 IEEE World AI IoT Congress (AIIoT);2023-06-07