Incremental hierarchical memory size estimation for steering of loop transformations-Reference-Cited by-同舟云学术

Incremental hierarchical memory size estimation for steering of loop transformations

Published:2007-09 Issue:4 Volume:12 Page:50
ISSN:1084-4309
Container-title:ACM Transactions on Design Automation of Electronic Systems
language:en
Short-container-title:ACM Trans. Des. Autom. Electron. Syst.

Author:

Hu Q.¹,Kjeldsberg P. G.¹,Vandecappelle A.²,Palkovic M.²,Catthoor F.²

Affiliation:

1. Norwegian University of Science and Technology, Trondheim, Norway

2. IMEC, Leuven, Belgium

Abstract

Modern embedded multimedia and telecommunications systems need to store and access huge amounts of data. This becomes a critical factor for the overall energy consumption, area, and performance of the systems. Loop transformations are essential to improve the data access locality and regularity in order to optimally design or utilize a memory hierarchy. However, due to abstract high-level cost functions, current loop transformation steering techniques do not take the memory platform sufficiently into account. They usually also result in only one final transformation solution. On the other hand, the loop transformation search space for real-life applications is huge, especially if the memory platform is still not fully fixed. Use of existing loop transformation techniques will therefore typically lead to suboptimal end-products. It is critical to find all interesting loop transformation instances. This can only be achieved by performing an evaluation of the effect of later design stages at the early loop transformation stage. This article presents a fast incremental hierarchical memory-size requirement estimation technique. It estimates the influence of any given sequence of loop transformation instances on the mapping of application data onto a hierarchical memory platform. As the exact memory platform instantiation is often not yet defined at this high-level design stage, a platform-independent estimation is introduced with a Pareto curve output for each loop transformation instance. Comparison among the Pareto curves helps the designer, or a steering tool, to find all interesting loop transformation instances that might later lead to low-power data mapping for any of the many possible memory hierarchy instances. Initially, the source code is used as input for estimation. However, performing the estimation repeatedly from the source code is too slow for large search space exploration. An incremental approach, based on local updating of the previous result, is therefore used to handle sequences of different loop transformations. Experiments show that the initial approach takes a few seconds, which is two orders of magnitude faster than state-of-the-art solutions but still too costly to be performed interactively many times. The incremental approach typically takes just a few milliseconds, which is another two orders of magnitude faster than the initial approach. This huge speedup allows us for the first time to handle real-life industrial-size applications and get realistic feedback during loop transformation exploration.

Publisher

Association for Computing Machinery (ACM)

Subject

Electrical and Electronic Engineering,Computer Graphics and Computer-Aided Design,Computer Science Applications

Link

https://dl.acm.org/doi/pdf/10.1145/1278349.1278363

Reference41 articles.

1. Compiler transformations for high-performance computing

2. Background memory area estimation for multidimensional signal processing systems

3. Banerjee U. 1993. Loop Transformation for Restructuring Compilers: The Foundations. Kluwer Academic Boston MA. Banerjee U. 1993. Loop Transformation for Restructuring Compilers: The Foundations. Kluwer Academic Boston MA.

4. A study of replacement algorithms for a virtual-storage computer

5. Increasing energy efficiency of embedded systems by application-specific memory hierarchy generation

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A methodology correlating code optimizations with data memory accesses, execution time and energy consumption;The Journal of Supercomputing;2019-05-13

2. Integrating Memory Optimization with Mapping Algorithms for Multi-Processors System-on-Chip;ACM Transactions on Embedded Computing Systems;2012-09

3. Design of Image Processing Embedded Systems Using Multidimensional Data Flow;EMBED SYST;2011

4. Design Space Exploration for Efficient Data Intensive Computing on SoCs;Handbook of Data Intensive Computing;2011

5. Constructing Application-Specific Memory Hierarchies on FPGAs;Transactions on High-Performance Embedded Architectures and Compilers III;2011