A Text Classification Model via Multi-Level Semantic Features-Reference-Cited by-同舟云学术

A Text Classification Model via Multi-Level Semantic Features

Published:2022-09-17 Issue:9 Volume:14 Page:1938
ISSN:2073-8994
Container-title:Symmetry
language:en
Short-container-title:Symmetry

Author:

Mao Keji^ORCID,Xu Jinyu,Yao Xingda,Qiu Jiefan,Chi Kaikai,Dai Guanglin

Abstract

Text classification is a major task of NLP (Natural Language Processing) and has been the focus of attention for years. News classification as a branch of text classification is characterized by complex structure, large amounts of information and long text length, which in turn leads to a decrease in the accuracy of classification. To improve the classification accuracy of Chinese news texts, we present a text classification model based on multi-level semantic features. First, we add the category correlation coefficient to TF-IDF (Term Frequency-Inverse Document Frequency) and the frequency concentration coefficient to CHI (Chi-Square), and extract the keyword semantic features with the improved algorithm. Then, we extract local semantic features with TextCNN with symmetric-channel and global semantic information from a BiLSTM with attention. Finally, we fuse the three semantic features for the prediction of text categories. The results of experiments on THUCNews, LTNews and MCNews show that our presented method is highly accurate, with 98.01%, 90.95% and 94.24% accuracy, respectively. With model parameters two magnitudes smaller than Bert, the improvements relative to the baseline Bert+FC are 1.27%, 1.2%, and 2.81%, respectively.

Funder

Basic Public Welfare Research Project of Zhejiang Province

National Natural Science Foundation of China

Publisher

MDPI AG

Subject

Physics and Astronomy (miscellaneous),General Mathematics,Chemistry (miscellaneous),Computer Science (miscellaneous)

Link

https://www.mdpi.com/2073-8994/14/9/1938/pdf

Reference40 articles.

1. A Sentiment Classification Method of Web Social Media Based on Multidimensional and Multilevel Modeling

2. TopicBERT: A Topic-Enhanced Neural Language Model Fine-Tuned for Sentiment Classification

3. SaTYa: Trusted Bi-LSTM-Based Fake News Classification Scheme for Smart Community

4. An Evolutionary Fake News Detection Method for COVID-19 Pandemic Information

5. An LSTM&Topic-CNN Model for Classification of Online Chinese Medical Questions

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Efficient Agricultural Question Classification With a BERT-Enhanced DPCNN Model;IEEE Access;2024

2. Explainable Machine Learning Models for Swahili News Classification;Proceedings of the 2023 7th International Conference on Natural Language Processing and Information Retrieval;2023-12-15

3. Classification of Text on Social Media Data Using the TF-IDF Approach, Word2Vec and Transfer Learning;2023 10th International Conference on Electrical Engineering, Computer Science and Informatics (EECSI);2023-09-20

4. From Optimal Control to Mean Field Optimal Transport via Stochastic Neural Networks;Symmetry;2023-09-08

5. Adaptive Dimensional Gaussian Mutation of PSO-Optimized Convolutional Neural Network Hyperparameters;Applied Sciences;2023-03-27