A Data-Driven Model for Automated Chinese Word Segmentation and POS Tagging-Reference-Cited by-同舟云学术

A Data-Driven Model for Automated Chinese Word Segmentation and POS Tagging

Published:2022-09-16 Issue: Volume:2022 Page:1-10
ISSN:1687-5273
Container-title:Computational Intelligence and Neuroscience
language:en
Short-container-title:Computational Intelligence and Neuroscience

Author:

Xu Qing¹^ORCID,Wang Zhiyou²^ORCID

Affiliation:

1. Changsha University of Science and Technology, Changsha, Hunan 410000, China

2. School of Electronic Communication and Electrical Engineering, Changsha University, Changsha, Hunan 410000, China

Abstract

Chinese natural language processing tasks often require the solution of Chinese word segmentation and POS tagging problems. Traditional Chinese word segmentation and POS tagging methods mainly use simple matching algorithms based on lexicons and rules. The simple matching or statistical analysis requires manual word segmentation followed by POS tagging, which leads to the inability to meet the practical requirements for label prediction accuracy. With the continuous development of deep learning technology, data-driven machine learning models provide new opportunities for automated Chinese word segmentation and POS tagging. Therefore, a data-driven automated Chinese word segmentation and POS tagging model is proposed in order to address the above problems. Firstly, the main idea and overall framework of the proposed automated model are outlined, and the tagging strategy and neural network language model used are described. Secondly, two main optimisations are made on the input side of the model: (1) the use of word2Vec for the representation of text features, thus representing the text as a distributed word vector; and (2) the use of an improved AlexNet for efficient encoding of long-range word, and the addition of an attention mechanism to the model. Finally, on the output side, an additional auxiliary loss function was designed to optimise the Chinese text based on its frequency. The experimental results show that the proposed model can significantly improve the accuracy and operational efficiency of Chinese word segmentation and POS tagging compared with other existing models, thus verifying its effectiveness and advancement.

Funder

Research on the Application and Practice of Medical Esp Teaching

Publisher

Hindawi Limited

Subject

General Mathematics,General Medicine,General Neuroscience,General Computer Science

Link

http://downloads.hindawi.com/journals/cin/2022/7622392.pdf

Reference44 articles.

1. Natural language processing: an introduction;P. M. Nadkarni;Journal of the American Medical Informatics Association,2011

2. Natural language processing;A. Chopra;International Journal of Technology Enhancements and Emerging Engineering Research,2013

3. A Primer on Neural Network Models for Natural Language Processing

4. Attention in Natural Language Processing

5. Introduction to Arabic Natural Language Processing

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Assessment of Digital Twin Virtual-Real Connected Interaction Based on NLP;2023 2nd International Conference on Artificial Intelligence and Intelligent Information Processing (AIIIP);2023-10-27