PerTract: Model Extraction and Specification of Big Data Systems for Performance Prediction by the Example of Apache Spark and Hadoop-Reference-Cited by-同舟云学术

PerTract: Model Extraction and Specification of Big Data Systems for Performance Prediction by the Example of Apache Spark and Hadoop

Published:2019-08-09 Issue:3 Volume:3 Page:47
ISSN:2504-2289
Container-title:Big Data and Cognitive Computing
language:en
Short-container-title:BDCC

Author:

Kroß Johannes^ORCID,Krcmar Helmut^ORCID

Abstract

Evaluating and predicting the performance of big data applications are required to efficiently size capacities and manage operations. Gaining profound insights into the system architecture, dependencies of components, resource demands, and configurations cause difficulties to engineers. To address these challenges, this paper presents an approach to automatically extract and transform system specifications to predict the performance of applications. It consists of three components. First, a system-and tool-agnostic domain-specific language (DSL) allows the modeling of performance-relevant factors of big data applications, computing resources, and data workload. Second, DSL instances are automatically extracted from monitored measurements of Apache Spark and Apache Hadoop (i.e., YARN and HDFS) systems. Third, these instances are transformed to model- and simulation-based performance evaluation tools to allow predictions. By adapting DSL instances, our approach enables engineers to predict the performance of applications for different scenarios such as changing data input and resources. We evaluate our approach by predicting the performance of linear regression and random forest applications of the HiBench benchmark suite. Simulation results of adjusted DSL instances compared to measurement results show accurate predictions errors below 15% based upon averages for response times and resource utilization.

Publisher

MDPI AG

Subject

Artificial Intelligence,Computer Science Applications,Information Systems,Management Information Systems

Link

https://www.mdpi.com/2504-2289/3/3/47/pdf

Reference44 articles.

1. Big Data

2. Performance Management Work

3. Quantitative Evaluation of Model-Driven Performance Analysis and Simulation of Component-Based Architectures

4. Performance-Oriented DevOps: A Research Agenda;Brunnert,2015

Cited by 9 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Application of Natural Language Processing and Genetic Algorithm to Fine-Tune Hyperparameters of Classifiers for Economic Activities Analysis;Big Data and Cognitive Computing;2024-06-13

2. Modeling and Simulating Stream Processing Platforms;2023 Winter Simulation Conference (WSC);2023-12-10

3. Evaluating Task-Level CPU Efficiency for Distributed Stream Processing Systems;Big Data and Cognitive Computing;2023-03-10

4. DICE simulation: a tool for software performance assessment at the design stage;Automated Software Engineering;2022-03-28

5. Accurate Performance Predictions with Component-Based Models of Data Streaming Applications;Software Architecture;2022