A Benchmark for Morphological Segmentation in Uyghur and Kazakh-Reference-Cited by-同舟云学术

A Benchmark for Morphological Segmentation in Uyghur and Kazakh

Published:2024-06-21 Issue:13 Volume:14 Page:5369
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Abudouwaili Gulinigeer¹²,Ruzmamat Sirajahmat¹²,Abiderexiti Kahaerjiang¹²,Wu Binghong¹²,Wumaier Aishan¹²^ORCID

Affiliation:

1. School of Computer Science and Technology, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China

2. Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China

Abstract

Morphological segmentation and stemming are foundational tasks in natural language processing. They have become effective ways to alleviate data sparsity in agglutinative languages because of the nature of agglutinative language word formation. Uyghur and Kazakh, as typical agglutinative languages, have made significant progress in morphological segmentation and stemming in recent years. However, the evaluation metrics used in previous work are character-level based, which may not comprehensively reflect the performance of models in morphological segmentation or stemming. Moreover, existing methods avoid manual feature extraction, but the model’s ability to learn features is inadequate in complex scenarios, and the correlation between different features has not been considered. Consequently, these models lack representation in complex contexts, affecting their effective generalization in practical scenarios. To address these issues, this paper redefines the morphological-level evaluation metrics: F1-score and accuracy (ACC) for morphological segmentation and stemming tasks. In addition, two models are proposed for morpheme segmentation and stem extraction tasks: supervised model and unsupervised model. The supervised model learns character and contextual features simultaneously, then feature embeddings are input into a Transformer encoder to study the correlation between character and context embeddings. The last layer of the model uses a CRF or softmax layer to determine morphological boundaries. In unsupervised learning, an encoder–decoder structure introduces n-gram correlation assumptions and masked attention mechanisms, enhancing the correlation between characters within n-grams and reducing the impact of characters outside n-grams on boundaries. Finally, comprehensive comparative analyses of the performance of different models are conducted from various points of view. Experimental results demonstrate that: (1) The proposed evaluation method effectively reflects the differences in morphological segmentation and stemming for Uyghur and Kazakh; (2) Learning different features and their correlation can enhance the model’s generalization ability in complex contexts. The proposed models achieve state-of-the-art performance on Uyghur and Kazakh datasets.

Funder

National Natural Science Foundation of China

Natural Science Foundation of Xinjiang Province

Publisher

MDPI AG

Link

https://www.mdpi.com/2076-3417/14/13/5369/pdf

Reference45 articles.

1. Sorokin, A. (2019, January 2). Convolutional neural networks for low-resource morpheme segmentation: Baseline or state-of-the-art?. Proceedings of the 16th Workshop on Computational Research in Phonetics, Phonology, and Morphology, Florence, Italy.

2. AgglutiFiT: Efficient Low-Resource Agglutinative Language Model Fine-Tuning;Li;IEEE Access,2020

3. Text-to-Speech for Low-Resource Agglutinative Language With Morphology-Aware Language Model Pre-Training;Liu;IEEE/ACM Trans. Audio Speech Lang. Process.,2024

4. Parhat, S., Sattar, M., Hamdulla, A., and Kadir, A. (2023). Uyghur–Kazakh–Kirghiz Text Keyword Extraction Based on Morpheme Segmentation. Information, 14.

5. Pan, Y., Li, X., Yang, Y., and Dong, R. (2020, January 5–10). Multi-Task Neural Model for Agglutinative Language Translation. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop, Online.