A new Chinese text clustering algorithm based on WRD and improved K-means-Reference-Cited by-同舟云学术

A new Chinese text clustering algorithm based on WRD and improved K-means

Published:2023-07-20 Issue:4 Volume:27 Page:1205-1220
ISSN:1088-467X
Container-title:Intelligent Data Analysis
language:
Short-container-title:IDA

Author:

Cui Zicai,Zhong Bocheng,Bai Chen

Abstract

Text clustering has been widely used in data mining, document management, search engines, and other fields. The K-means algorithm is a representative algorithm of text clustering. However, traditional K-means algorithm often uses Euclidean distance or cosine distance to measure the similarity between texts, which is not effective in face of high-dimensional data and cannot retain enough semantic information. In response to the above problems, we combine word rotator’s distance with the K-means algorithm, and propose the WRDK-means algorithm, which use word rotator’s distance to calculate the similarity between texts and preserve more text features. Furthermore, we define a new cluster center initialization method that improves cluster instability during random initial cluster center selection. And, to solve the problem of inconsistent length between texts, we propose a new iterative approximation method of cluster centers. We selected three suitable datasets and five evaluation indicators to verify the feasibility of the proposed algorithm. Among them, the RI value of our algorithm exceeds 90%. And for Marco_F1, our scheme was about 37.77%, 23.2%, 13.06% and 20.12% better than other four methods, respectively.

Publisher

IOS Press

Subject

Artificial Intelligence,Computer Vision and Pattern Recognition,Theoretical Computer Science

Reference18 articles.

1. Review on the Research of K-means Clustering Algorithm in Big Data

2. W. Ling, C. Dyer, A.W. Black et al., Two/too simple adaptations of word2vec for syntax problems, in: Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2015, pp. 1299–1304.

3. M.Q. James, Some methods for classification and analysis of multivariate observations, Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability 1(14) (1967).

4. J. Lee and J.H. Lee, K-means clustering based SVM ensemble methods for imbalanced data problem, in: 2014 Joint 7th International Conference on Soft Computing and Intelligent Systems (SCIS) and 15th International Symposium on Advanced Intelligent Systems (ISIS), 2014, pp. 614–617.

5. Improved fast partitional clustering algorithm for text clustering;Bejos;Journal of Intelligent and Fuzzy Systems,2020

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Improved K-means clustering-genetic backpropagation modeling for online state-of-charge estimation of lithium-ion batteries adaptive to low-temperature conditions;Journal of Energy Storage;2024-10

2. Research on Topic Identification Algorithm Based on Semantic Clustering: A Case Study of the Metaverse Research;Proceedings of the 2023 7th International Conference on Electronic Information Technology and Computer Engineering;2023-10-20