Evaluating Controlled Memory Request Injection for Efficient Bandwidth Utilization and Predictable Execution in Heterogeneous SoCs-Reference-Cited by-同舟云学术

Evaluating Controlled Memory Request Injection for Efficient Bandwidth Utilization and Predictable Execution in Heterogeneous SoCs

Published:2022-12-13 Issue:1 Volume:22 Page:1-25
ISSN:1539-9087
Container-title:ACM Transactions on Embedded Computing Systems
language:en
Short-container-title:ACM Trans. Embed. Comput. Syst.

Author:

Brilli Gianluca¹^ORCID,Cavicchioli Roberto²^ORCID,Solieri Marco³^ORCID,Valente Paolo³^ORCID,Marongiu Andrea³^ORCID

Affiliation:

1. Department of ‘Ingegneria Enzo Ferrari’, University of Modena and Reggio Emilia, Modena, Europe

2. Department of Sciences and Methods for Engineering, University of Modena and Reggio Emilia, Europe

3. Department of Physics, Informatics and Mathematics, University of Modena and Reggio Emilia, Modena, Italy

Abstract

High-performance embedded platforms are increasingly adopting heterogeneous systems-on-chip (HeSoC) that couple multi-core CPUs with accelerators such as GPU, FPGA, or AI engines. Adopting HeSoCs in the context of real-time workloads is not immediately possible, though, as contention on shared resources like the memory hierarchy—and in particular the main memory (DRAM)—causes unpredictable latency increase. To tackle this problem, both the research community and certification authorities mandate (i) that accesses from parallel threads to the shared system resources (typically, main memory) happen in a mutually exclusive manner by design, or (ii) that per-thread bandwidth regulation is enforced. Such arbitration schemes provide timing guarantees, but make poor use of the memory bandwidth available in a modern HeSoC. Controlled Memory Request Injection (CMRI) is a recently-proposed bandwidth limitation concept that builds on top of a mutually-exclusive schedule but still allows the threads currently not entitled to access memory to use as much of the unused bandwidth as possible without losing the timing guarantee. CMRI has been discussed in the context of a multi-core CPU, but the same principle applies also to a more complex system such as an HeSoC. In this article, we introduce two CMRI schemes suitable for HeSoCs: Voluntary Throttling via code refactoring and Bandwidth Regulation via dynamic throttling. We extensively characterize a proof-of-concept incarnation of both schemes on two HeSoCs: an NVIDIA Tegra TX2 and a Xilinx UltraScale+, highlighting the benefits and the costs of CMRI for synthetic workloads that model worst-case DRAM access. We also test the effectiveness of CMRI with real benchmarks, studying the effect of interference among the host CPU and the accelerators.

Funder

ECSEL JU projects COMP4DRONES

AI4CSM

Publisher

Association for Computing Machinery (ACM)

Subject

Hardware and Architecture,Software

Link

https://dl.acm.org/doi/pdf/10.1145/3548773

Reference37 articles.

1. [n.d.]. Solving Multicore Interference for Safety-Critical Applications. Retrieved from https://www.ghs.com/download/whitepapers/GHS_multicore_interference.pdf.

2. On the tailoring of CAST-32A certification guidance to real COTS multicore architectures

3. Managing Heterogeneous Resources in HPC Systems

4. Schedulability analysis of global memory-predictable scheduling

5. Time-predictable execution of multithreaded applications on multicore systems