Towards a holistic approach to auto-parallelization-Reference-Cited by-同舟云学术

Towards a holistic approach to auto-parallelization

Published:2009-05-28 Issue:6 Volume:44 Page:177-187
ISSN:0362-1340
Container-title:ACM SIGPLAN Notices
language:en
Short-container-title:SIGPLAN Not.

Author:

Tournavitis Georgios¹,Wang Zheng¹,Franke Björn¹,O'Boyle Michael F.P.¹

Affiliation:

1. University of Edinburgh, Edinburgh, United Kingdom

Abstract

Compiler-based auto-parallelization is a much studied area, yet has still not found wide-spread application. This is largely due to the poor exploitation of application parallelism, subsequently resulting in performance levels far below those which a skilled expert programmer could achieve. We have identified two weaknesses in traditional parallelizing compilers and propose a novel, integrated approach, resulting in significant performance improvements of the generated parallel code. Using profile-driven parallelism detection we overcome the limitations of static analysis, enabling us to identify more application parallelism and only rely on the user for final approval. In addition, we replace the traditional target-specific and inflexible mapping heuristics with a machine-learning based prediction mechanism, resulting in better mapping decisions while providing more scope for adaptation to different target architectures. We have evaluated our parallelization strategy against the NAS and SPEC OMP benchmarks and two different multi-core platforms (dual quad-core Intel Xeon SMP and dual-socket QS20 Cell blade). We demonstrate that our approach not only yields significant improvements when compared with state-of-the-art parallelizing compilers, but comes close to and sometimes exceeds the performance of manually parallelized codes. On average, our methodology achieves 96% of the performance of the hand-tuned OpenMP NAS and SPEC parallel benchmarks on the Intel Xeon platform and gains a significant speedup for the IBM Cell platform, demonstrating the potential of profile-guided and machine-learning based parallelization for complex multi-core platforms.

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Graphics and Computer-Aided Design,Software

Link

https://dl.acm.org/doi/pdf/10.1145/1543135.1542496

Reference45 articles.

1. Future Microprocessors and Off-Chip SOP Interconnect

2. The parallel execution of DO loops

3. Interprocedural dependence analysis and parallelization

4. Maximizing parallelism and minimizing synchronization with affine partitions

Cited by 67 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. IDaTPA: importance degree based thread partitioning approach in thread level speculation;Discover Computing;2024-06-19

2. Towards Green AI: Current Status and Future Research;2024 Electronics Goes Green 2024+ (EGG);2024-06-18

3. A new thread-level speculative automatic parallelization model and library based on duplicate code execution;The Journal of Supercomputing;2024-03-11

4. Revealing Compiler Heuristics Through Automated Discovery and Optimization;2024 IEEE/ACM International Symposium on Code Generation and Optimization (CGO);2024-03-02

5. Investigating the superiority of Intel oneAPI IFX compiler on Intel CPUs using different optimization levels: A case study on a CFD system;2023 4th IEEE Global Conference for Advancement in Technology (GCAT);2023-10-06