Semantic-Based Classification of Long Texts on Higher Education in China-Reference-Cited by-同舟云学术

Semantic-Based Classification of Long Texts on Higher Education in China

Published:2021-10-20 Issue: Volume:2021 Page:1-8
ISSN:1607-887X
Container-title:Discrete Dynamics in Nature and Society
language:en
Short-container-title:Discrete Dynamics in Nature and Society

Author:

Li Chun¹^ORCID,Fei Yanying²

Affiliation:

1. School of Marxism, Dalian University of Technology, Dalian 116023, China

2. Faculty of Humanities and Social Sciences, Dalian University of Technology, Dalian 116086, China

Abstract

The development level of higher education (HE) is an important indicator of the development level and development potential of a country. The HE-related document is the mirror to reflect the develop process of the HE. The research of high education (HE) has been developing rapidly in China, resulting in a huge number of texts, such as relevant policies, speech drafts, and yearbooks. The traditional manual classification of HE texts is inefficient and unable to deal with the huge number of HE texts. Besides, the effect of direct classification is rather poor because HE texts tend to be long and exist as an imbalanced dataset. To solve these problems, this paper improves the convolutional neural network (CNN) into the HE-CNN classification model for HE texts. Firstly, Chinese HE policies, speech drafts, and yearbooks (1979–2020) were downloaded from the official website of Chinese Ministry of Education. In total, 463 files were collected and divided into four classes, namely, definition, task, method, and effect evaluation. To handle the huge number of HE texts, the Twitter-latent Dirichlet allocation (LDA) topic model was employed to extract word frequency and critical information, such as age and author, enhancing the training effect of CNN. To address the dataset imbalance problem, CNN parameters were optimized repeatedly through comparative experiments, which further improve the training effect. Finally, the proposed HE-CNN model was found more effective and accurate than other classification models.

Funder

Key Project of Liaoning Provincial Law Society

Publisher

Hindawi Limited

Subject

Modeling and Simulation

Link

http://downloads.hindawi.com/journals/ddns/2021/9237713.pdf

Reference29 articles.

1. Analysis of the Questioning Characteristics of Elementary Science Gifted Education Teaching Materials using the Sternberg’s View of Successful Intelligence: Focused on Semantic Network Analysis

2. Query Rewriting and Semantic Annotation in Semantic-Based Image Retrieval under Heterogeneous Ontologies of Big Data

3. Semantic Web Mining for Analyzing Retail Environment Using Word2Vec and CNN-FK

4. Research on Text Sentiment Analysis Based on Neural Network and Ensemble Learning

5. A Novel Spatio-Temporal Violence Classification Framework Based on Material Derivative and LSTM Neural Network

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Retracted: Semantic-Based Classification of Long Texts on Higher Education in China;Discrete Dynamics in Nature and Society;2024-01-24