KCB-FLAT: Enhancing Chinese Named Entity Recognition with Syntactic Information and Boundary Smoothing Techniques-Reference-Cited by-同舟云学术

KCB-FLAT: Enhancing Chinese Named Entity Recognition with Syntactic Information and Boundary Smoothing Techniques

Published:2024-08-30 Issue:17 Volume:12 Page:2714
ISSN:2227-7390
Container-title:Mathematics
language:en
Short-container-title:Mathematics

Author:

Deng Zhenrong¹²,Huang Zheng¹,Wei Shiwei³^ORCID,Zhang Jinglin¹

Affiliation:

1. Guangxi Key Laboratory of Images and Graphics Intelligent Processing, Guilin University of Electronic Technology, Guilin 541004, China

2. Nanning Research Institute, Guilin University of Electronic Technology, Guilin 541004, China

3. School of Computer Science and Engineering, Guilin University of Aerospace Technology, Guilin 541004, China

Abstract

Named entity recognition (NER) is a fundamental task in Natural Language Processing (NLP). During the training process, NER models suffer from over-confidence, and especially for the Chinese NER task, it involves word segmentation and introduces erroneous entity boundary segmentation, exacerbating over-confidence and reducing the model’s overall performance. These issues limit further enhancement of NER models. To tackle these problems, we proposes a new model named KCB-FLAT, designed to enhance Chinese NER performance by integrating enriched semantic information with the word-Boundary Smoothing technique. Particularly, we first extract various types of syntactic data and utilize a network named Key-Value Memory Network, based on syntactic information to functionalize this, integrating it through an attention mechanism to generate syntactic feature embeddings for Chinese characters. Subsequently, we employed an encoder named Cross-Transformer to thoroughly combine syntactic and lexical information to address the entity boundary segmentation errors caused by lexical information. Finally, we introduce a Boundary Smoothing module, combined with a regularity-conscious function, to capture the internal regularity of per entity, reducing the model’s overconfidence in entity probabilities through smoothing. Experimental results demonstrate that the proposed model achieves exceptional performance on the MSRA, Resume, Weibo, and self-built ZJ datasets, as verified by the F1 score.

Funder

Guangxi Science and Technology Project

National Natural Science Foundation of China

Guangxi Key Laboratory of Image and Graphic Intelligent Processing Project

Innovation Project of GUET Graduate Education

Publisher

MDPI AG

Link

https://www.mdpi.com/2227-7390/12/17/2714/pdf

Reference50 articles.

1. Yin, D., Cheng, S., Pan, B., Qiao, Y., Zhao, W., and Wang, D. (2022). Chinese Named Entity Recognition Based on Knowledge Based Question Answering System. Appl. Sci., 12.

2. Bose, P., Srinivasan, S., Sleeman, W.C., Palta, J., Kapoor, R., and Ghosh, P. (2021). A Survey on Recent Named Entity Recognition and Relationship Extraction Techniques on Clinical Texts. Appl. Sci., 11.

3. Chen, S., Pei, Y., Ke, Z., and Silamu, W. (2021). Low-Resource Named Entity Recognition via the Pre-Training Model. Symmetry, 13.

4. Ahmad, P.N., Shah, A.M., and Lee, K. (2023). A Review on Electronic Health Record Text-Mining for Biomedical Name Entity Recognition in Healthcare Domain. Healthcare, 11.

5. Huang, C., Wang, Y., Yu, Y., Hao, Y., Liu, Y., and Zhao, X. (2022). Chinese Named Entity Recognition of Geological News Based on BERT Model. Appl. Sci., 12.