Comparison of Pre-trained vs Custom-trained Word Embedding Models for Word Sense Disambiguation-Reference-Cited by-同舟云学术

Comparison of Pre-trained vs Custom-trained Word Embedding Models for Word Sense Disambiguation

Published:2023-11-01 Issue:1 Volume:12 Page:e31084
ISSN:2255-2863
Container-title:ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal
language:
Short-container-title:ADCAIJ

Author:

Farhat Ullah Muhammad,Saeed Ali,Hussain Naveed

Abstract

The prime objective of word sense disambiguation (WSD) is to develop such machines that can automatically recognize the actual meaning (sense) of ambiguous words in a sentence. WSD can improve various NLP and HCI challenges. Researchers explored a wide variety of methods to resolve this issue of sense ambiguity. However, majorly, their focus was on English and some other well-reputed languages. Urdu with more than 300 million users and a large amount of electronic text available on the web is still unexplored. In recent years, for a variety of Natural Language Processing tasks, word embedding methods have proven extremely successful. This study evaluates, compares, and applies a variety of word embedding approaches to Urdu Word embedding (both Lexical Sample and All-Words), including pre-trained (Word2Vec, Glove, and FastText) as well as custom-trained (Word2Vec, Glove, and FastText trained on the Ur-Mono corpus). Two benchmark corpora are used for the evaluation in this study: (1) the UAW-WSD-18 corpus and (2) the ULS-WSD-18 corpus. For Urdu All-Words WSD tasks, top results have been achieved (Accuracy=60.07 and F1=0.45) using pre-trained FastText. For the Lexical Sample, WSD has been achieved (Accuracy=70.93 and F1=0.60) using custom-trained GloVe word embedding method.

Publisher

Ediciones Universidad de Salamanca

Subject

Information Systems,Computer Science Applications,Computer Networks and Communications,Artificial Intelligence

Reference40 articles.

1. Abid, M., A. H., Jawad, A., and Abdul, S., 2018. Urdu word sense disambiguation using machine learning approach. Cluster Computing 21(1), 515–522. 10.1007/s10586-017-0918-0

2. Ali, M., N., and Tan, G., and Hussain, A., 2018. Bidirectional recurrent neural network approach for Arabic named entity recognition. Future Internet. 10(12), 123. 10.3390/fi10120123

3. Ali, S., Nawab, R. M. A., Mark, S., and Paul, R., 2019. A word sense disambiguation corpus for Urdu. Language Resources and Evaluation. 53: 397–418. 10.1007/s10579-018-9438-7

4. Ali, S., Rao, M. A. N., Mark, S., and Paul, R., 2019. A Sense Annotated Corpus for All-Words Urdu Word Sense Disambiguation. ACM Transactions on Asian and Low-Resource Language Information Processing, 18(4), 1–14. 10.1145/3314940

5. Archana, K., and DK, L., 2020. Word2vec’s Distributed Word Representation for Hindi Word Sense Disambiguation. In International Conference on Distributed Computing and Internet Technology, 325–335. 10.1007/978-3-030-36987-3_21