Keyword Data Analysis Using Generative Models Based on Statistics and Machine Learning Algorithms-Reference-Cited by-同舟云学术

Keyword Data Analysis Using Generative Models Based on Statistics and Machine Learning Algorithms

Published:2024-02-19 Issue:4 Volume:13 Page:798
ISSN:2079-9292
Container-title:Electronics
language:en
Short-container-title:Electronics

Author:

Jun Sunghae¹^ORCID

Affiliation:

1. Department of Data Science, Cheongju University, Cheongju 28503, Chungbuk, Republic of Korea

Abstract

For text big data analysis, we preprocessed text data and constructed a document–keyword matrix. The elements of this matrix represent the frequencies of keywords occurring in a document. The matrix has a zero-inflation problem because many elements are zero values. Also, in the process of preprocessing, the data size of the document–keyword matrix is reduced. However, various machine learning algorithms require a large amount of data, so to solve the problems of data shortage and zero inflation, we propose the use of generative models based on statistics and machine learning. In our experimental tests, we compared the performance of the models using simulation and practical data sets. Thus, we verified the validity and contribution of our research for keyword data analysis.

Publisher

MDPI AG

Link

https://www.mdpi.com/2079-9292/13/4/798/pdf

Reference43 articles.

1. Jun, S. (2023). Zero-Inflated Text Data Analysis using Generative Adversarial Networks and Statistical Modeling. Computers, 12.

2. General-use unsupervised keyword extraction model for keyword analysis;Shin;Expert Syst. Appl.,2023

3. Digital business foresight: Keyword-based analysis and CorEx topic modeling;Bzhalava;Futures,2024

4. Julia, S., and Robinson, D. (2017). Text Mining with R, O’Reilly.

5. Feinerer, I., and Hornik, K. (2023). Package ‘tm’ Version 0.7-11, Text Mining Package, CRAN of R Project, R Foundation for Statistical Computing.