Reaching bandwidth saturation using transparent injection parallelization-Reference-Cited by-同舟云学术

Reaching bandwidth saturation using transparent injection parallelization

Published:2016-11-09 Issue:5 Volume:31 Page:405-421
ISSN:1094-3420
Container-title:The International Journal of High Performance Computing Applications
language:en
Short-container-title:The International Journal of High Performance Computing Applications

Author:

Chaimov Nicholas¹,Ibrahim Khaled Z²,Williams Samuel²,Iancu Costin²

Affiliation:

1. University of Oregon, OR, USA

2. Lawrence Berkeley National Laboratory, CA, USA

Abstract

Although logically available, applications may not exploit enough instantaneous communication concurrency to maximize network utilization on HPC systems. This is exacerbated in hybrid programming models that combine single program multiple data with OpenMP or CUDA. We present the design of a “multi-threaded” runtime able to transparently increase the instantaneous network concurrency and to provide near saturation bandwidth, independent of the application configuration and dynamic behavior. The runtime offloads communication requests from application level tasks to multiple communication servers. The servers use system specific performance models to attain network saturation. Our techniques alleviate the need for spatial and temporal application level message concurrency optimizations. Experimental results show improved message throughput and bandwidth by as much as 150% for 4 KB messages on InfiniBand and by as much as 120% for 4 KB messages on Cray Aries. For more complex operations such as all-to-all collectives, we observe as much as 30% speedup. This translates into 23% speedup on 12,288 cores for a NAS FT implemented using FFTW. We observe as much as 76% speedup on 1500 cores for an already optimized UPC+OpenMP geometric multigrid application using hybrid parallelism. For the geometric multigrid GPU implementation, we observe as much as 44% speedup on 512 GPUs.

Funder

Advanced Scientific Computing Research

Office of Science

Publisher

SAGE Publications

Subject

Hardware and Architecture,Theoretical Computer Science,Software

Link

http://journals.sagepub.com/doi/pdf/10.1177/1094342016672720

Reference27 articles.

1. The NAS parallel benchmarks---summary and preliminary results

2. MPI ON MILLIONS OF CORES

3. Hybrid PGAS runtime support for multicore nodes

4. MPI versus MPI+OpenMP on the IBM SP for the NAS Benchmarks

5. X10

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Special issue on programming models and applications for multicores and manycores;The International Journal of High Performance Computing Applications;2017-08-23