A Tradeoff Analysis of FPGAs, GPUs, and Multicores for Sliding-Window Applications-Reference-Cited by-同舟云学术

A Tradeoff Analysis of FPGAs, GPUs, and Multicores for Sliding-Window Applications

Published:2015-03-06 Issue:1 Volume:8 Page:1-24
ISSN:1936-7406
Container-title:ACM Transactions on Reconfigurable Technology and Systems
language:en
Short-container-title:ACM Trans. Reconfigurable Technol. Syst.

Author:

Cooke Patrick¹,Fowers Jeremy¹,Brown Greg¹,Stitt Greg¹

Affiliation:

1. University of Florida, Gainesville, USA

Abstract

The increasing usage of hardware accelerators such as Field-Programmable Gate Arrays (FPGAs) and Graphics Processing Units (GPUs) has significantly increased application design complexity. Such complexity results from a larger design space created by numerous combinations of accelerators, algorithms, and hw/sw partitions. Exploration of this increased design space is critical due to widely varying performance and energy consumption for each accelerator when used for different application domains and different use cases. To address this problem, numerous studies have evaluated specific applications across different architectures. In this article, we analyze an important domain of applications, referred to as sliding-window applications , implemented on FPGAs, GPUs, and multicore CPUs. For each device, we present optimization strategies and analyze use cases where each device is most effective. The results show that, for large input sizes, FPGAs can achieve speedups of up to 5.6× and 58× compared to GPUs and multicore CPUs, respectively, while also using up to an order of magnitude less energy. For small input sizes and applications with frequency-domain algorithms, GPUs generally provide the best performance and energy.

Funder

National Science Foundation

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/2659000

Reference31 articles.

1. Altera. 2013. Altera’s User-Customizable ARM-Based SoC. (2013). Retrieved from http://www.altera.com/literature/br/br-soc-fpga.pdf. Altera. 2013. Altera’s User-Customizable ARM-Based SoC. (2013). Retrieved from http://www.altera.com/literature/br/br-soc-fpga.pdf.

2. Performance comparison of FPGA, GPU and CPU in image processing

3. A Configurable Processor Synthesis System

4. AMD Fusion APU: Llano

5. Real-Time Optical Flow Calculations on FPGA and GPU Architectures: A Comparison Study

Cited by 19 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Real-time Operational Load Monitoring of a Composite Aerostructure Using FPGA-based Computing System;Bulletin of the Polish Academy of Sciences Technical Sciences;2023-11-02

2. Design and Simulation of a Median Filter for a CubeSat Image Processing Application Using an FPGA Architecture;ITM Web of Conferences;2022

3. ShuntFlowPlus: An Efficient and Scalable Dataflow Accelerator Architecture for Stream Applications;ACM J EMERG TECH COM;2021

4. cuZ-Checker: A GPU-Based Ultra-Fast Assessment System for Lossy Compressions;2021 IEEE International Conference on Cluster Computing (CLUSTER);2021-09

5. FPGA Implementations of Algorithms for Preprocessing of High Frame Rate and High Resolution Image Streams in Real Time;Annals of Emerging Technologies in Computing;2021-04-01