Affiliation:
1. College of Information Engineering, Northwest A&F University, 3 Taicheng Road, Yangling, Xianyang 712100, China
2. Research Center of Information Technology, Beijing Academy of Agriculture and Forestry Sciences, Beijing 100097, China
Abstract
Feature selection is crucial in classification tasks as it helps to extract relevant information while reducing redundancy. This paper presents a novel method that considers both instance and label correlation. By employing the least squares method, we calculate the linear relationship between each feature and the target variable, resulting in correlation coefficients. Features with high correlation coefficients are selected. Compared to traditional methods, our approach offers two advantages. Firstly, it effectively selects features highly correlated with the target variable from a large feature set, reducing data dimensionality and improving analysis and modeling efficiency. Secondly, our method considers label correlation between features, enhancing the accuracy of selected features and subsequent model performance. Experimental results on three datasets demonstrate the effectiveness of our method in selecting features with high correlation coefficients, leading to superior model performance. Notably, our approach achieves a minimum accuracy improvement of 3.2% for the advanced classifier, lightGBM, surpassing other feature selection methods. In summary, our proposed method, based on instance and label correlation, presents a suitable solution for classification problems.
Reference24 articles.
1. Sidey-Gibbons, J.A.M., and Sidey-Gibbons, C.J. (2019). Machine learning in medicine: A practical introduction. BMC Med. Res. Methodol., 19.
2. Machine learning regression and classification methods for fog events prediction;Ghimire;Atmos. Res.,2022
3. Shlens, J. (2014). A Tutorial on Principal Component Analysis. arXiv.
4. Meesad, P., Boonrawd, P., and Nuipian, V. (2011, January 28–29). A Chi-Square-Test for Word Importance Differentiation in Text Classification. Proceedings of the International Conference on Information and Electronics Engineering, Bangkok, Thailand.
5. Exploring feature selection and classification methods for predicting heart disease;Spencer;Digit. Health,2020