Machine learning based software effort estimation using development-centric features for crowdsourcing platform-Reference-Cited by-同舟云学术

Machine learning based software effort estimation using development-centric features for crowdsourcing platform

Published:2023-11-30 Issue: Volume: Page:1-31
ISSN:1088-467X
Container-title:Intelligent Data Analysis
language:
Short-container-title:IDA

Author:

Yasmin Anum¹,Haider Wasi¹,Daud Ali²³,Banjar Ameen³

Affiliation:

1. Department of Computer and Software Engcineering, College of Electrical and Mechanical Engineering, National University of Sciences and Technology (NUST), Islamabad, Pakistan

2. Abu Dhabi School of Management, Abu Dhabi, United Arab Emirates

3. Department of Information systems and Technology, College of Computer Science and Engineering, University of Jeddah, Jeddah, Saudi Arabia

Abstract

Crowd-Sourced software development (CSSD) is getting a good deal of attention from the software and research community in recent times. One of the key challenges faced by CSSD platforms is the task selection mechanism which in practice, contains no intelligent scheme. Rather, rule-of-thumb or intuition strategies are employed, leading to biasness and subjectivity. Effort considerations on crowdsourced tasks can offer good foundation for task selection criteria but are not much investigated. Software development effort estimation (SDEE) is quite prevalent domain in software engineering but only investigated for in-house development. For open-sourced or crowdsourced platforms, it is rarely explored. Moreover, Machine learning (ML) techniques are overpowering SDEE with a claim to provide more accurate estimation results. This work aims to conjoin ML-based SDEE to analyze development effort measures on CSSD platform. The purpose is to discover development-oriented features for crowdsourced tasks and analyze performance of ML techniques to find best estimation model on CSSD dataset. TopCoder is selected as target CSSD platform for the study. TopCoder’s development tasks data with development-centric features are extracted, leading to statistical, regression and correlation analysis to justify features’ significance. For effort estimation, 10 ML families with 2 respective techniques are applied to get broader aspect of estimation. Five performance metrices (MSE, RMSE, MMRE, MdMRE, Pred (25) and Welch’s statistical test are incorporated to judge the worth of effort estimation model’s performance. Data analysis results show that selected features of TopCoder pertain reasonable model significance, regression, and correlation measures. Findings of ML effort estimation depicted that best results for TopCoder dataset can be acquired by linear, non-linear regression and SVM family models. To conclude, the study identified the most relevant development features for CSSD platform, confirmed by in-depth data analysis. This reflects careful selection of effort estimation features to offer good basis of accurate ML estimate.

Publisher

IOS Press

Subject

Artificial Intelligence,Computer Vision and Pattern Recognition,Theoretical Computer Science

Reference74 articles.

1. T. Alelyani, K. Mao and Y. Yang, Context-centric pricing: early pricing models for software crowdsourcing tasks, in: Proceedings of the 13th International Conference on Predictive Models and Data Analytics in Software Engineering, 2017, pp. 63–72.

2. T. Alelyani and Y. Yang, Software crowdsourcing reliability: an empirical study on developers behavior, in: Proceedings of the 2nd International Workshop on Software Analytics, 2016, pp. 36–42.

3. A. Ali and C. Gravino, Using bio-inspired features selection algorithms in software effort estimation: a systematic literature review, in: 2019 45th Euromicro Conference on Software Engineering and Advanced Applications (SEAA), IEEE, 2019, pp. 220–227.

4. Software development effort estimation using classical and fuzzy analogy: A cross-validation comparative study;Amazal;International Journal of Computational Intelligence Applications,2014

5. Empirical analysis on productivity prediction and locality for use case points method;Azzeh;Software Quality Journal,2021