Efficient parallelization of SPH algorithm on modern multi-core CPUs and massively parallel GPUs-Reference-Cited by-同舟云学术

Efficient parallelization of SPH algorithm on modern multi-core CPUs and massively parallel GPUs

Published:2021-07-10 Issue:06 Volume:12 Page:2150054
ISSN:1793-9623
Container-title:International Journal of Modeling, Simulation, and Scientific Computing
language:en
Short-container-title:Int. J. Model. Simul. Sci. Comput.

Author:

Jagtap Pravin¹,Nasre Rupesh²,Sanapala V. S.³,Patnaik B. S. V.¹^ORCID

Affiliation:

1. Department of Applied Mechanics, Indian Institute of Technology Madras, Chennai 600036, India

2. Department of Computer Science & Engineering, Indian Institute of Technology Madras, Chennai 600036, India

3. Indira Gandhi Centre for Atomic Research, Homi Bhabha National Institute, Kalpakkam 603 102, India

Abstract

Smoothed Particle Hydrodynamics (SPH) is fast emerging as a practically useful computational simulation tool for a wide variety of engineering problems. SPH is also gaining popularity as the back bone for fast and realistic animations in graphics and video games. The Lagrangian and mesh-free nature of the method facilitates fast and accurate simulation of material deformation, interface capture, etc. Typically, particle-based methods would necessitate particle search and locate algorithms to be implemented efficiently, as continuous creation of neighbor particle lists is a computationally expensive step. Hence, it is advantageous to implement SPH, on modern multi-core platforms with the help of High-Performance Computing (HPC) tools. In this work, the computational performance of an SPH algorithm is assessed on multi-core Central Processing Unit (CPU) as well as massively parallel General Purpose Graphical Processing Units (GP-GPU). Parallelizing SPH faces several challenges such as, scalability of the neighbor search process, force calculations, minimizing thread divergence, achieving coalesced memory access patterns, balancing workload, ensuring optimum use of computational resources, etc. While addressing some of these challenges, detailed analysis of performance metrics such as speedup, global load efficiency, global store efficiency, warp execution efficiency, occupancy, etc. is evaluated. The OpenMP and Compute Unified Device Architecture[Formula: see text] parallel programming models have been used for parallel computing on Intel Xeon[Formula: see text] E5-[Formula: see text] multi-core CPU and NVIDIA Quadro M[Formula: see text] and NVIDIA Tesla p[Formula: see text] massively parallel GPU architectures. Standard benchmark problems from the Computational Fluid Dynamics (CFD) literature are chosen for the validation. The key concern of how to identify a suitable architecture for mesh-less methods which essentially require heavy workload of neighbor search and evaluation of local force fields from neighbor interactions is addressed.

Publisher

World Scientific Pub Co Pte Ltd

Subject

Computer Science Applications,Modelling and Simulation

Link

https://www.worldscientific.com/doi/pdf/10.1142/S1793962321500549

Reference41 articles.

1. A new algorithm for solving some mechanical problems

2. A NEW BENCHMARK QUALITY SOLUTION FOR THE BUOYANCY-DRIVEN CAVITY BY DISCRETE SINGULAR CONVOLUTION

3. Wavelets generated by using discrete singular convolution kernels

4. High order matched interface and boundary method for elliptic equations with discontinuous coefficients and singular sources

5. Smoothed particle hydrodynamics: theory and application to non-spherical stars

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Establishment and validation of a viscous-potential coupled and graphics processing unit accelerated numerical tank based on smoothed particle hydrodynamics and high-order spectral methods;Physics of Fluids;2023-10-01

2. Evaluation of Separation Efficiency in a Tubular Bowl Centrifuge;Chemical Engineering & Technology;2022-01-12