EMOGI-Reference-Cited by-同舟云学术

EMOGI

Published:2020-10 Issue:2 Volume:14 Page:114-127
ISSN:2150-8097
Container-title:Proceedings of the VLDB Endowment
language:en
Short-container-title:Proc. VLDB Endow.

Author:

Min Seung Won¹,Mailthody Vikram Sharma¹,Qureshi Zaid¹,Xiong Jinjun²,Ebrahimi Eiman³,Hwu Wen-mei¹

Affiliation:

1. University of Illinois at Urbana-Champaign

2. IBM T.J. Watson Research Center Yorktown Heights

3. NVIDIA

Abstract

Modern analytics and recommendation systems are increasingly based on graph data that capture the relations between entities being analyzed. Practical graphs come in huge sizes, offer massive parallelism, and are stored in sparse-matrix formats such as compressed sparse row (CSR). To exploit the massive parallelism, developers are increasingly interested in using GPUs for graph traversal. However, due to their sizes, graphs often do not fit into the GPU memory. Prior works have either used input data pre-processing/partitioning or unified virtual memory (UVM) to migrate chunks of data from the host memory to the GPU memory. However, the large, multi-dimensional, and sparse nature of graph data presents a major challenge to these schemes and results in significant amplification of data movement and reduced effective data throughput. In this work, we propose EMOGI, an alternative approach to traverse graphs that do not fit in GPU memory using direct cache-line-sized access to data stored in host memory. This paper addresses the open question of whether a sufficiently large number of overlapping cache-line-sized accesses can be sustained to 1) tolerate the long latency to host memory, 2) fully utilize the available bandwidth, and 3) achieve favorable execution performance. We analyze the data access patterns of several graph traversal applications in GPU over PCIe using an FPGA to understand the cause of poor external bandwidth utilization. By carefully coalescing and aligning external memory requests, we show that we can minimize the number of PCIe transactions and nearly fully utilize the PCIe bandwidth with direct cache-line accesses to the host memory. EMOGI achieves 2.60X speedup on average compared to the optimized UVM implementations in various graph traversal applications. We also show that EMOGI scales better than a UVM-based solution when the system uses higher bandwidth interconnects such as PCIe 4.0.

Publisher

VLDB Endowment

Subject

General Earth and Planetary Sciences,Water Science and Technology,Geography, Planning and Development

Link

https://dl.acm.org/doi/pdf/10.14778/3425879.3425883

Cited by 24 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Shared Virtual Memory: Its Design and Performance Implications for Diverse Applications;Proceedings of the 38th ACM International Conference on Supercomputing;2024-05-30

2. MultiEM: Efficient and Effective Unsupervised Multi-Table Entity Matching;2024 IEEE 40th International Conference on Data Engineering (ICDE);2024-05-13

3. GMT: GPU Orchestrated Memory Tiering for the Big Data Era;Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3;2024-04-27

4. Graph Processing Scheme Using GPU With Value-Driven Differential Scheduling;IEEE Access;2024

5. HongTu: Scalable Full-Graph GNN Training on Multiple GPUs;Proceedings of the ACM on Management of Data;2023-12-08