Comparison of Data Mining Classification Algorithms Determining the Default Risk-Reference-Cited by-同舟云学术

Comparison of Data Mining Classification Algorithms Determining the Default Risk

Published:2019-02-03 Issue: Volume:2019 Page:1-8
ISSN:1058-9244
Container-title:Scientific Programming
language:en
Short-container-title:Scientific Programming

Author:

Çığşar Begüm¹,Ünal Deniz¹^ORCID

Affiliation:

1. Cukurova University, Faculty of Arts and Sciences, Department of Statistics, Adana, Turkey

Abstract

Big data and its analysis have become a widespread practice in recent times, applicable to multiple industries. Data mining is a technique that is based on statistical applications. This method extracts previously undetermined data items from large quantities of data. The banking and insurance industries use data mining analysis to detect fraud, offer the appropriate credit or insurance solutions to customers, and better understand customer demands. This study aims to identify data mining classification algorithms and use them to predict default risks, avoid possible payment difficulties, and reduce potential problems in extending credit. The data for this study, which contains demographic and socioeconomic characteristics of individuals, were obtained from the Turkish Statistical Institute 2015 survey. Six classification algorithms—Naive Bayes, Bayesian networks, J48, random forest, multilayer perceptron, and logistic regression—were applied to the dataset using WEKA 3.9 data mining software. These algorithms were compared considering the root mean error squares, receiver operating characteristic area, accuracy, precision, F-measure, and recall statistical criteria. The best algorithm—logistic regression—was obtained and applied to the real dataset to determine the attributes causing the default risk by using odds ratios. The socioeconomic and demographic characteristics of the individuals were examined, and based on the odds ratio values, the results of which individuals and characteristics were more likely to default, were reached. These results are not only beneficial to the literature but also have a significant influence in the financial industry in terms of the ability to predict customers’ default risk.

Funder

Cukurova University Scientific Research Fund

Publisher

Hindawi Limited

Subject

Computer Science Applications,Software

Link

http://downloads.hindawi.com/journals/sp/2019/8706505.pdf

Reference25 articles.

1. Comparative Analysis of Classification Algorithms on Different Datasets using WEKA

2. Artificial Intelligence and Data Mining: Algorithms and Applications

Cited by 37 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. مقارنة خوارزميات تنقيب البيانات;المجلة الليبية العالمية;2024-03-19

2. Congestion Control Prediction Model for 5G Environment Based on Supervised and Unsupervised Machine Learning Approach;IEEE Access;2024

3. Machine Learning-based BGP Traffic Prediction;2023 IEEE 22nd International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom);2023-11-01

4. Assessment of open-pit captive limestone mining areas using sentinel-2 imagery with spectral indices and machine learning algorithms;International Journal of Knowledge-based and Intelligent Engineering Systems;2023-10-13

5. A recent review on optimisation methods applied to credit scoring models;Journal of Economics, Finance and Administrative Science;2023-06-05