Coordinated CTA Combination and Bandwidth Partitioning for GPU Concurrent Kernel Execution-Reference-Cited by-同舟云学术

Coordinated CTA Combination and Bandwidth Partitioning for GPU Concurrent Kernel Execution

Published:2019-08-20 Issue:3 Volume:16 Page:1-27
ISSN:1544-3566
Container-title:ACM Transactions on Architecture and Code Optimization
language:en
Short-container-title:ACM Trans. Archit. Code Optim.

Author:

Lin Zhen¹,Dai Hongwen¹,Mantor Michael²,Zhou Huiyang¹

Affiliation:

1. North Carolina State University, Raleigh, NC, USA

2. Advanced Micro Devices, Orlando, FL, USA

Abstract

Contemporary GPUs support multiple kernels to run concurrently on the same streaming multiprocessors (SMs). Recent studies have demonstrated that such concurrent kernel execution (CKE) improves both resource utilization and computational throughput. Most of the prior works focus on partitioning the GPU resources at the cooperative thread array (CTA) level or the warp scheduler level to improve CKE. However, significant performance slowdown and unfairness are observed when latency-sensitive kernels co-run with bandwidth-intensive ones. The reason is that bandwidth over-subscription from bandwidth-intensive kernels leads to much aggravated memory access latency, which is highly detrimental to latency-sensitive kernels. Even among bandwidth-intensive kernels, more intensive kernels may unfairly consume much higher bandwidth than less-intensive ones. In this article, we first make a case that such problems cannot be sufficiently solved by managing CTA combinations alone and reveal the fundamental reasons. Then, we propose a coordinated approach for CTA combination and bandwidth partitioning. Our approach dynamically detects co-running kernels as latency sensitive or bandwidth intensive. As both the DRAM bandwidth and L2-to-L1 Network-on-Chip (NoC) bandwidth can be the critical resource, our approach partitions both bandwidth resources coordinately along with selecting proper CTA combinations. The key objective is to allocate more CTA resources for latency-sensitive kernels and more NoC/DRAM bandwidth resources to NoC-/DRAM-intensive kernels. We achieve it using a variation of dominant resource fairness (DRF). Compared with two state-of-the-art CKE optimization schemes, SMK [52] and WS [55], our approach improves the average harmonic speedup by 78% and 39%, respectively. Even compared to the best possible CTA combinations, which are obtained from an exhaustive search among all possible CTA combinations, our approach improves the harmonic speedup by up to 51% and 11% on average.

Funder

National science foundation

Publisher

Association for Computing Machinery (ACM)

Subject

Hardware and Architecture,Information Systems,Software

Link

https://dl.acm.org/doi/pdf/10.1145/3326124

Reference59 articles.

Cited by 8 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. GPU-Based Algorithms for Processing the k Nearest-Neighbor Query on Spatial Data Using Partitioning and Concurrent Kernel Execution;International Journal of Parallel Programming;2023-07-21

2. LATOA: Load-Aware Task Offloading and Adoption in GPU;Proceedings of the 15th Workshop on General Purpose Processing Using GPU;2023-02-25

3. BARM: A Batch-Aware Resource Manager for Boosting Multiple Neural Networks Inference on GPUs With Memory Oversubscription;IEEE Transactions on Parallel and Distributed Systems;2022-12-01

4. GPUPool;Proceedings of the International Conference on Parallel Architectures and Compilation Techniques;2022-10-08

5. A Survey of GPU Multitasking Methods Supported by Hardware Architecture;IEEE Transactions on Parallel and Distributed Systems;2022-06-01