An Improved Corpus-Based NLP Method for Facilitating Keyword Extraction: An Example of the COVID-19 Vaccine Hesitancy Corpus-Reference-Cited by-同舟云学术

An Improved Corpus-Based NLP Method for Facilitating Keyword Extraction: An Example of the COVID-19 Vaccine Hesitancy Corpus

Published:2023-02-13 Issue:4 Volume:15 Page:3402
ISSN:2071-1050
Container-title:Sustainability
language:en
Short-container-title:Sustainability

Author:

Chen Liang-Ching¹²^ORCID

Affiliation:

1. Department of Foreign Languages, R.O.C. Military Academy, Kaohsiung 830, Taiwan

2. Institute of Education, National Sun Yat-sen University, Kaohsiung 804, Taiwan

Abstract

In the current COVID-19 post-pandemic era, COVID-19 vaccine hesitancy is hindering the herd immunity generated by widespread vaccination. It is critical to identify the factors that may cause COVID-19 vaccine hesitancy, enabling the relevant authorities to propose appropriate interventions for mitigating such a phenomenon. Keyword extraction, a sub-field of natural language processing (NLP) applications, plays a vital role in modern medical informatics. When traditional corpus-based NLP methods are used to conduct keyword extraction, they only consider a word’s log-likelihood value to determine whether it is a keyword, which leaves room for concerns about the efficiency and accuracy of this keyword extraction technique. These concerns include the fact that the method is unable to (1) optimize the keyword list by the machine-based approach, (2) effectively evaluate the keyword’s importance level, and (3) integrate the variables to conduct data clustering. Thus, to address the aforementioned issues, this study integrated a machine-based word removal technique, the i10-index, and the importance–performance analysis (IPA) technique to develop an improved corpus-based NLP method for facilitating keyword extraction. The top 200 most-cited Science Citation Index (SCI) research articles discussing COVID-19 vaccine hesitancy were adopted as the target corpus for verification. The results showed that the keywords of Quadrant I (n = 98) reached the highest lexical coverage (9.81%), indicating that the proposed method successfully identified and extracted the most important keywords from the target corpus, thus achieving more domain-oriented and accurate keyword extraction results.

Publisher

MDPI AG

Subject

Management, Monitoring, Policy and Law,Renewable Energy, Sustainability and the Environment,Geography, Planning and Development,Building and Construction

Link

https://www.mdpi.com/2071-1050/15/4/3402/pdf

Reference62 articles.

1. Natural language processing enabling COVID-19 predictive analytics to support data-driven patient advising and pooled testing;Meystre;J. Am. Med Inf. Assoc.,2021

2. A survey on different dimensions for graphical keyword extraction techniques issues and challenges;Garg;Artif. Intell. Rev.,2021

3. Mao, K.J., Xu, J.Y., Yao, X.D., Qiu, J.F., Chi, K.K., and Dai, G.L. (2022). A text classification model via multi-level semantic features. Symmetry, 14.

4. Trappey, A.J.C., Liang, C.P., and Lin, H.J. (2022). Using machine learning language models to generate innovation knowledge graphs for patent mining. Appl. Sci., 12.

5. Accurate methods for the statistics of surprise and coincidence;Dunning;Comput. Linguist.,1993

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. An extended TF-IDF method for improving keyword extraction in traditional corpus-based research: An example of a climate change corpus;Data & Knowledge Engineering;2024-09

2. An entropy-based corpus method for improving keyword extraction: An example of sustainability corpus;Engineering Applications of Artificial Intelligence;2024-07

3. A Short-Text Similarity Model Combining Semantic and Syntactic Information;Electronics;2023-07-18

4. University Student Dropout Prediction Using Pretrained Language Models;Applied Sciences;2023-06-13