Text Feature Extraction for Public English Vocabulary Based on Wavelet Transform-Reference-Cited by-同舟云学术

Text Feature Extraction for Public English Vocabulary Based on Wavelet Transform

Published:2022-06-11 Issue: Volume:2022 Page:1-10
ISSN:1748-6718
Container-title:Computational and Mathematical Methods in Medicine
language:en
Short-container-title:Computational and Mathematical Methods in Medicine

Author:

Ye Di¹^ORCID,Shi Xiaojing¹

Affiliation:

1. School of Foreign Studies, Tangshan Normal University, Tangshan, Hebei 063000, China

Abstract

Text interpretation of public English vocabulary is a critical task in the subject of natural language processing, which uses technology to allow humans and computers to communicate effectively using natural language. Text feature extraction is one of the most fundamental and crucial elements in allowing computers to effectively grasp and read text. This paper proposes a text feature extraction method based on wavelet analysis that performs fast discrete wavelet transform and inverse discrete wavelet transform on the feature vectors under the traditional TF-IDF vector space model to address the problem of low feature differentiation of high-dimensional data in text feature extraction. In particular, due to the design of the Mallat algorithm, there is frequency aliasing in the signal decomposition process. This phenomenon is a problem that cannot be ignored when using wavelet analysis for feature extraction. Therefore, this paper proposes an improved inverse discrete wavelet transform method, in which the signal is decomposed by Mallat algorithm to obtain wavelet coefficients at each scale and then reconstructed to the required wavelet space coefficients according to the reconstruction method, and the reconstructed coefficients are used to analyze the signal at that scale instead of the wavelet coefficients obtained at the corresponding scale. Experiments on the public English vocabulary dataset reveal that the wavelet transform-based strategy suggested in this research outperforms existing feature extraction methods while maintaining greater classification accuracy while reducing the dimensionality of the TF-IDF vector space model.

Publisher

Hindawi Limited

Subject

Applied Mathematics,General Immunology and Microbiology,General Biochemistry, Genetics and Molecular Biology,Modeling and Simulation,General Medicine

Link

http://downloads.hindawi.com/journals/cmmm/2022/7125242.pdf

Reference22 articles.

1. Feature selection using hybrid poor and rich optimization algorithm for text classification

2. Feature selection based on feature interactions with application to text categorization

3. WAVELET LOSSY COMPRESSION OF RANDOM DATA

4. Overview of Compressed Sensing: Sensing Model, Reconstruction Algorithm, and Its Applications

5. Towards a Model to Improve Boolean Knowledge Mapping by Using Text Mining and Its Applications

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Retracted: Text Feature Extraction for Public English Vocabulary Based on Wavelet Transform;Computational and Mathematical Methods in Medicine;2023-09-20