OptimML: Joint Control of Inference Latency and Server Power Consumption for ML Performance Optimization-Reference-Cited by-同舟云学术

OptimML: Joint Control of Inference Latency and Server Power Consumption for ML Performance Optimization

Published:2024-05-07 Issue: Volume: Page:
ISSN:1556-4665
Container-title:ACM Transactions on Autonomous and Adaptive Systems
language:en
Short-container-title:ACM Trans. Auton. Adapt. Syst.

Author:

Chen Guoyu¹^ORCID,Wang Xiaorui¹^ORCID

Affiliation:

1. The Ohio State University, USA

Abstract

Power capping is an important technique for high-density servers to safely oversubscribe the power infrastructure in a data center. However, power capping is commonly accomplished by dynamically lowering the server processors’ frequency levels, which can result in degraded application performance. For servers that run important machine learning (ML) applications with Service-Level Objective (SLO) requirements, inference performance such as recognition accuracy must be optimized within a certain latency constraint, which demands high server performance. In order to achieve the best inference accuracy under the desired latency and server power constraints, this paper proposes OptimML, a multi-input-multi-output (MIMO) control framework that jointly controls both inference latency and server power consumption, by flexibly adjusting the machine learning model size (and so its required computing resources) when server frequency needs to be lowered for power capping. Our results on a hardware testbed with widely adopted ML framework (including PyTorch, TensorFlow, and MXNet) show that OptimML achieves higher inference accuracy compared with several well-designed baselines, while respecting both latency and power constraints. Furthermore, an adaptive control scheme with online model switching and estimation is designed to achieve analytic assurance of control accuracy and system stability, even in the face of significant workload/hardware variations.

Publisher

Association for Computing Machinery (ACM)

Link

https://dl.acm.org/doi/pdf/10.1145/3661825

Reference56 articles.

1. Mohamed S Abdelfattah, Łukasz Dudziak, Thomas Chau, Royson Lee, Hyeji Kim, and Nicholas D Lane. 2020. Best of both worlds: Automl codesign of a cnn and its hardware accelerator. In 2020 57th ACM/IEEE Design Automation Conference (DAC). IEEE, 1–6.

2. Shekhar Borkar and Andrew A. Chien. 2011. The Future of Microprocessors. Communications of the ACM, Vol. 54 No. 5, Pages 67-77 (2011).

3. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 33. Curran Associates, Inc., 1877–1901.

4. Model slicing for supporting complex analytics with elastic inference cost and resource constraints

5. Ming Chen, Xiaorui Wang, and Xue Li. 2011. Coordinating Processor and Main Memory for Efficient Server Power Control. In Proceedings of the 25th International Conference on Supercomputing (ICS).