The Impact of Data Quality on Software Testing Effort Prediction-Reference-Cited by-同舟云学术

The Impact of Data Quality on Software Testing Effort Prediction

Published:2023-03-31 Issue:7 Volume:12 Page:1656
ISSN:2079-9292
Container-title:Electronics
language:en
Short-container-title:Electronics

Author:

Radliński Łukasz¹^ORCID

Affiliation:

1. Faculty of Computer Science and Information Technology, West Pomeranian University of Technology in Szczecin, ul. Żołnierska 49, 71-210 Szczecin, Poland

Abstract

Background: This paper investigates the impact of data quality on the performance of models predicting effort on software testing. Data quality was reflected by training data filtering strategies (data variants) covering combinations of Data Quality Rating, UFP Rating, and a threshold of valid cases. Methods: The experiment used the ISBSG dataset and 16 machine learning models. A process of three-fold cross-validation repeated 20 times was used to train and evaluate each model with each data variant. Model performance was assessed using absolute errors of prediction. A ‘win–tie–loss’ procedure, based on the Wilcoxon signed-rank test, was applied to identify the best models and data variants. Results: Most models, especially the most accurate, performed the best on a complete dataset, even though it contained cases with low data ratings. The detailed results include the rankings of the following: (1) models for particular data variants, (2) data variants for particular models, and (3) the best-performing combinations of models and data variants. Conclusions: Arbitrary and restrictive data selection to only projects with Data Quality Rating and UFP Rating of ‘A’ or ‘B’, commonly used in the literature, does not seem justified. It is recommended not to exclude cases with low data ratings to achieve better accuracy of most predictive models for testing effort prediction.

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Computer Networks and Communications,Hardware and Architecture,Signal Processing,Control and Systems Engineering

Link

https://www.mdpi.com/2079-9292/12/7/1656/pdf

Reference60 articles.

1. Systematic literature review of machine learning based software development effort estimation models;Wen;Inf. Softw. Technol.,2012

2. A Systematic Review of Software Development Cost Estimation Studies;Jorgensen;IEEE Trans. Softw. Eng.,2007

3. Ali, A., and Gravino, C. (2019). A systematic literature review of software effort prediction using machine learning methods. J. Softw. Evol. Process., 31.

4. Software development effort estimation: A systematic mapping study;Farias;IET Softw.,2020

5. Software effort estimation accuracy prediction of machine learning techniques: A systematic performance evaluation;Mahmood;Softw. Pract. Exp.,2022

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. The METRIC-framework for assessing data quality for trustworthy AI in medicine: a systematic review;npj Digital Medicine;2024-08-03

2. Gradient Boosting Optimized Through Differential Evolution for Predicting the Testing Effort of Software Projects;IEEE Access;2023