Oracle-guided scheduling for controlling granularity in implicitly parallel languages-Reference-Cited by-同舟云学术

Oracle-guided scheduling for controlling granularity in implicitly parallel languages

Published:2016 Issue: Volume:26 Page:
ISSN:0956-7968
Container-title:Journal of Functional Programming
language:en
Short-container-title:J. Funct. Prog.

Author:

ACAR UMUT A.,CHARGUÉRAUD ARTHUR,RAINEY MIKE

Abstract

AbstractA classic problem in parallel computing is determining whether to execute a thread in parallel or sequentially. If small threads are executed in parallel, the overheads due to thread creation can overwhelm the benefits of parallelism, resulting in suboptimal efficiency and performance. If large threads are executed sequentially, processors may spin idle, resulting again in sub-optimal efficiency and performance. This “granularity problem” is especially important in implicitly parallel languages, where the programmer expresses all potential for parallelism, leaving it to the system to exploit parallelism by creating threads as necessary. Although this problem has been identified as an important problem, it is not well understood—broadly applicable solutions remain elusive. In this paper, we propose techniques for automatically controlling granularity in implicitly parallel programming languages to achieve parallel efficiency and performance. To this end, we first extend a classic result, Brent's theorem (a.k.a. the work-time principle) to include thread-creation overheads. Using a cost semantics for a general-purpose language in the style of lambda calculus with parallel tuples, we then present a precise accounting of thread-creation overheads and bound their impact on efficiency and performance. To reduce such overheads, we propose an oracle-guided semantics by using estimates of the sizes of parallel threads. We show that, if the oracle provides accurate estimates in constant time, then the oracle-guided semantics reduces the thread-creation overheads for a reasonably large class of parallel computations. We describe how to approximate the oracle-guided semantics in practice by combining static and dynamic techniques. We require the programmer to provide the asymptotic complexity cost for each parallel thread and use runtime profiling to determine hardware-specific constant factors. We present an implementation of the proposed approach as an extension of the Manticore compiler for Parallel ML. Our empirical evaluation shows that our techniques can reduce thread-creation overheads, leading to good efficiency and performance.

Publisher

Cambridge University Press (CUP)

Subject

Software

Reference54 articles.

1. On the Problem of Distribution in Globular Star Clusters: (Plate 8.)

2. Low-cost process creation and dynamic partitioning in Qlisp

3. Lazy task creation: a technique for increasing the granularity of parallel programs

4. Compiling collection-oriented languages onto massively parallel computers

Cited by 13 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Efficient Parallel Functional Programming with Effects;Proceedings of the ACM on Programming Languages;2023-06-06

2. Responsive Parallelism with Synchronization;Proceedings of the ACM on Programming Languages;2023-06-06

3. Task parallel assembly language for uncompromising parallelism;Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation;2021-06-18

4. Provably space-efficient parallel functional programming;Proceedings of the ACM on Programming Languages;2021-01-04

5. Responsive parallelism with futures and state;Proceedings of the 41st ACM SIGPLAN Conference on Programming Language Design and Implementation;2020-06-06